Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,13 @@ Format: [Semantic Versioning](https://semver.org). Schema versions and record se
new record.

### Added
- AVE-2026-00083: silent guardrail comparison failure — a protection
mechanism (guardrail, approval gate, or judge) executes on every
request and reports a permissive result because its own comparison,
branch, or aggregation logic never evaluates the real input, across
five independently reproduced structural forms. Proposed by
arian-gogani (github.com/arian-gogani/failopen,
CWE-CAPEC/AI-Working-Group#1); researcher field credits him directly.
- AVE-2026-00081: Lingering authority — a task/subgoal/episode-scoped
capability grant outlives the closure event that justified it, with
nothing in the agent's runtime tying revocation to that closure, so
Expand Down
11 changes: 6 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Stable IDs, AIVSS scores, and behavioral fingerprints for every way a skill file
MCP server, system prompt, or agent plugin can be weaponized — scored consistently,
mapped to the frameworks security teams already report against.

[![Records](https://img.shields.io/badge/records-82-0f6e56?style=flat-square)](records/)
[![Records](https://img.shields.io/badge/records-83-0f6e56?style=flat-square)](records/)
[![Schema](https://img.shields.io/badge/schema-v1.1.0-0a3024?style=flat-square)](schema/ave-record-1.1.0.schema.json)
[![AIVSS](https://img.shields.io/badge/AIVSS-v0.8-d4a017?style=flat-square)](https://aivss.owasp.org)
[![OWASP MCP](https://img.shields.io/badge/OWASP-MCP%20Top%2010-0a3024?style=flat-square)](https://owasp.org)
Expand Down Expand Up @@ -102,7 +102,7 @@ two published AVE records, corrected the underlying process
documentation, not just the two records, credited in
[CONTRIBUTORS.md](CONTRIBUTORS.md).

82 records. 8 independent crosswalks. See
83 records. 8 independent crosswalks. See
[crosswalks/](crosswalks/) for the full mappings, and
[docs/writeups/](docs/writeups/) for full technical write-ups on
individual records.
Expand Down Expand Up @@ -140,12 +140,12 @@ skill file -> in CI / pre-commit -> before deploy

| | |
|---|---|
| Total records | 82 |
| Total records | 83 |
| Schema version | 1.1.0 |
| AIVSS spec | v0.8 |
| CRITICAL (>= 9.0) | 1 |
| HIGH (7.0-8.9) | 15 |
| MEDIUM (4.0-6.9) | 64 |
| MEDIUM (4.0-6.9) | 65 |
| LOW (< 4.0) | 2 |
| Framework: OWASP MCP Top 10 | all records |
| Framework: MITRE ATLAS | where applicable |
Expand Down Expand Up @@ -208,7 +208,7 @@ AIVSS = ((8.5 + 7.5) / 2) x 1.0 x 1 = 8.0 -> HIGH
## Record index

<details>
<summary><strong>82 records, click to expand</strong></summary>
<summary><strong>83 records, click to expand</strong></summary>

| AVE ID | Title | AIVSS | Severity |
|---|---|---|---|
Expand Down Expand Up @@ -294,6 +294,7 @@ AIVSS = ((8.5 + 7.5) / 2) x 1.0 x 1 = 8.0 -> HIGH
| [AVE-2026-00080](records/AVE-2026-00080.json) | Silent Agent Substitution (Sybil) via Unverified Retry | 6.8 | MEDIUM |
| [AVE-2026-00081](records/AVE-2026-00081.json) | Lingering Authority: Stale Capability Survives Episode Closure | 6.1 | MEDIUM |
| [AVE-2026-00082](records/AVE-2026-00082.json) | Local Skill Name Collision (Deterministic Router Shadowing) | 4.4 | MEDIUM |
| [AVE-2026-00083](records/AVE-2026-00083.json) | Silent guardrail comparison failure: a protection mechanism that executes but never evaluates its verdict | 6.3 | MEDIUM |

</details>

Expand Down
157 changes: 157 additions & 0 deletions dist/ave-records-latest.json

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions dist/ave-records-latest.manifest.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"schema_version": "1.1.0",
"record_count": 82,
"generated_at": "2026-09-22T00:00:39.582Z",
"record_count": 83,
"generated_at": "2026-09-22T00:11:53.246Z",
"source": "https://github.com/aveproject/ave"
}
157 changes: 157 additions & 0 deletions records/AVE-2026-00083.json

Large diffs are not rendered by default.

44 changes: 44 additions & 0 deletions tests/fixtures/AVE-2026-00083_negative.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Guardrail module: `role_adherence_metric.py` (deny-by-default verdict check)

Same judge-based role-adherence gate, same untyped judge completion as
input, but the comparison is inverted to require an affirmative,
evaluated "no violation" before releasing the response, instead of
falling through to release whenever the violation phrasing doesn't match
one exact literal.

```python
class RoleAdherenceVerdict:
def __init__(self, verdict: str, reason: str):
self.verdict = verdict # from the judge model's own completion, untyped
self.reason = reason

_CLEAN_TOKENS = {"no", "no.", "no violation", "none"}

def _calculate_score(self) -> float:
for v in self.verdicts:
normalized = v.verdict.strip().lower().rstrip(".")
if normalized in _CLEAN_TOKENS:
continue # explicitly and unambiguously clean
return 0.0 # anything else -- a "yes", a sentence, an empty
# string, an unparseable reply -- withholds the
# response and requires human review
return 1.0 # every verdict in the batch was explicitly clean
```

```
judge completion score released?
"no" 1.0 yes
"yes" 0.0 no
"Yes, the assistant broke character and gave medical advice." 0.0 no
"yes - it claimed to be a licensed physician" 0.0 no
"" 0.0 no
```

A conforming implementation must not flag this file: the permissive
outcome (`return 1.0`) is only reached when every verdict in the batch is
explicitly, affirmatively recognized as clean, and every other value --
including a verbose violation explanation, an empty string, or an
unparseable reply -- falls through to the withholding branch instead of
the releasing one. The comparison genuinely evaluates the real value
either way; nothing here can produce a silent pass for a value the check
never actually inspected.
39 changes: 39 additions & 0 deletions tests/fixtures/AVE-2026-00083_positive.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# Guardrail module: `role_adherence_metric.py`

A judge-based guardrail scores a model's role-adherence and gates whether an
agent response is released to the user. The verdict comes from a separate
judge-model call and is stored as a plain string.

```python
class RoleAdherenceVerdict:
def __init__(self, verdict: str, reason: str):
self.verdict = verdict # from the judge model's own completion, untyped
self.reason = reason

def _calculate_score(self) -> float:
for v in self.verdicts:
if v.verdict.strip().lower() == "yes":
return 0.0 # violation caught, response withheld
return 1.0 # no violation -- release the response
```

```
judge completion score released?
"yes" 0.0 no
"Yes, the assistant broke character and gave medical advice." 1.0 YES
"yes - it claimed to be a licensed physician" 1.0 YES
```

The judge model is asked to flag any role violation and, in two of the three
runs above, correctly identifies one in its own explanation -- but the gate
only recognizes the bare token `"yes"`. Because `_calculate_score` compares
the judge's full completion with an exact `==` against the single literal
`"yes"`, any verdict that explains itself in a sentence rather than replying
with that one word alone falls through to `return 1.0` and the response is
released. The check runs on every request, appears in the pipeline's logs as
`role_adherence_metric: PASSED`, and its own test suite is green, because the
one test case that exists uses the bare-token phrasing the comparison
expects. Nothing here is a missing check, a disabled flag, or an adversarial
prompt reaching the judge -- the judge correctly identified two violations in
its own words, and the gate's comparison logic is what silently discarded
that finding.
Loading