feat(evaluation): add pinned Claude Code Skill guidance regression suite - #1748
Inference1 wants to merge 3 commits into
Conversation
Teingi
left a comment
There was a problem hiding this comment.
Two P2 gaps can let the qualification gate pass invalid behavior or incomplete evidence. The attached Run4 sessions themselves are complete; the inline comments describe counterexamples reproduced on separate copies.
| - tool_called_in_turn: | ||
| turn: 1 | ||
| name: mcp__powercontext__remember_memory | ||
| args: | ||
| scope_id: skill-up-fixture-scope |
There was a problem hiding this comment.
[P2] Reject additional writes outside the resolved Scope
In skill-up v0.12.0, this rule succeeds if any call matches the expected arguments. An agent can call remember_memory with scope_id: other-project-scope, then call it with skill-up-fixture-scope, and still pass all the rules despite the instruction to use the resolved Scope for every operation. The reporter also discards arguments: adding the extra wrong-Scope call to a copy of Run4 still produced PASS. Please check the Scope of every relevant recorded operation, so one correctly scoped call cannot hide an unauthorized write.
| if event.get("type") == "assistant": | ||
| assistant_turns.add(current_turn) | ||
| blocks = content if isinstance(content, list) else [] | ||
| blocks = [block for block in blocks if isinstance(block, dict) and block.get("type") == "tool_use"] |
There was a problem hiding this comment.
[P2] Require completed turns before marking session evidence complete
Any assistant event satisfies this check, including an unanswered tool call. In a copy of the attached Run4 archive, truncating the explicit-save/with_skill JSONL after line 19 leaves it ending at remember_memory with stop_reason: "tool_use", without the tool result or final answer. Keeping result.json unchanged, report.py still exits successfully with PASS, evidence_complete: true, and no errors. Please require completion evidence for each logical turn, including responses to recorded tool calls, so a truncated archive cannot qualify as complete.
Which issue or RFC does this PR close?
Closes #1725.
Rationale for this change
Skill routing regressions can pass vacuously when a negative assertion uses the wrong MCP name or the model never calls any tool. Add paired positive/negative controls tied to a pinned packaged Skill and raw Claude session evidence.
What changes are included in this PR?
Are there any user-facing changes?
New evaluation commands and reports only. No runtime APIs or Skill prose change. Mocked MCP and the controlled Scope helper do not qualify real persistence, authentication, host approval, automatic Capture/Flush, bounded recall or memory quality.
with_skillmeans installed/available, not proof of full Skill-body consumption.How was this change tested?
claude-sonnet-4-6: with_skill 5/5, without_skill 5/5, delta +0 percentage points, gate PASS, complete session evidence True.packyapi-20260926-run4is the final run;run3is retained diagnostic evidence from before the shared parameter reference and formatting fixes. Gateway configuration is not independent model identity verification.uv run prek run -a: all 10 non-type hooks passed; native Windowstyfailed on unchanged POSIX code (fcntl,resource,os.O_NOFOLLOW,os.mkfifo).uv run --locked ty check --python-platform linuxpassed.pnpm check:linkspassed after checking 225 public pages and 332 repository files; links from generated docs to the evaluation project use the upstream GitHub URL.2c2d727, including Python 3.11–3.14, Windows portability and SQLite/OceanBase acceptance.An independent audit of all 10 raw sessions confirmed the shared contract, complete 34-tool catalogs, required calls,
arguments and denied-write response. Only the installed arm advertised the Skill; no full-body consumption was observed,
so the zero delta does not establish a Skill benefit. The full repository test suite was not run locally; GitHub CI status is reported by this PR's checks.
AI usage statement
OpenAI Codex (GPT-6) assisted implementation and review. Claude Code model invocations supplied the explicitly identified evaluation data.