Skip to content

feat(evaluation): add pinned Claude Code Skill guidance regression suite - #1748

Open
Inference1 wants to merge 3 commits into
oceanbase:masterfrom
Inference1:feat/1725-skill-guidance-evaluation
Open

Inference1 wants to merge 3 commits into
oceanbase:masterfrom
Inference1:feat/1725-skill-guidance-evaluation

Conversation

@Inference1

@Inference1 Inference1 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Which issue or RFC does this PR close?

Closes #1725.

Rationale for this change

Skill routing regressions can pass vacuously when a negative assertion uses the wrong MCP name or the model never calls any tool. Add paired positive/negative controls tied to a pinned packaged Skill and raw Claude session evidence.

What changes are included in this PR?

  • Add a skill-up v0.12.0 Claude Code suite for ordinary coding, explicit saves, empty search, candidate inspection and failed saves, with all 34 mocked operations and shared neutral parameter signatures.
  • Pin/vendor the unchanged packaged Skill; check drift and save exact input snapshots with each run.
  • Retain both arms and native JSON/JUnit/HTML plus raw JSONL; report exact tool names, case-level deltas and evidence completeness separately from behavior failures. Check every relevant call's Scope, matching tool results and a completed assistant response for every turn.
  • Add a credential-free CI validation job and complementary English/Chinese documentation.
  • Keep all 202 assertions while compacting repetitive YAML and removing general validation duplicated by the runner.

Are there any user-facing changes?

New evaluation commands and reports only. No runtime APIs or Skill prose change. Mocked MCP and the controlled Scope helper do not qualify real persistence, authentication, host approval, automatic Capture/Flush, bounded recall or memory quality. with_skill means installed/available, not proof of full Skill-body consumption.

How was this change tested?

  • skill-up validate/dry-run, pin/fixture validation: passed.
  • 26 unittest regressions, Ruff lint/format and local patch checks: passed.
  • Claude Code 2.1.283 through PackyAPI, requested claude-sonnet-4-6: with_skill 5/5, without_skill 5/5, delta +0 percentage points, gate PASS, complete session evidence True.
  • Complete evaluation evidence: native JSON/JUnit/HTML, raw session JSONL, exact input snapshots, gateway metadata and checksums. packyapi-20260926-run4 is the final run; run3 is retained diagnostic evidence from before the shared parameter reference and formatting fixes. Gateway configuration is not independent model identity verification.
  • Rechecked the retained Run4 sessions with the updated reporter: both arms remain 5/5, evidence complete. On separate copies, an extra wrong-Scope write and a transcript truncated at line 19 both passed the old reporter and fail the updated reporter. Original evidence hashes are unchanged; no new model run was needed because parsed cases, prompts, fixtures and Skill inputs are unchanged.
  • uv run prek run -a: all 10 non-type hooks passed; native Windows ty failed on unchanged POSIX code (fcntl, resource, os.O_NOFOLLOW, os.mkfifo). uv run --locked ty check --python-platform linux passed.
  • Website pnpm check:links passed after checking 225 public pages and 332 repository files; links from generated docs to the evaluation project use the upstream GitHub URL.
  • All 24 GitHub CI checks passed at 2c2d727, including Python 3.11–3.14, Windows portability and SQLite/OceanBase acceptance.

An independent audit of all 10 raw sessions confirmed the shared contract, complete 34-tool catalogs, required calls,
arguments and denied-write response. Only the installed arm advertised the Skill; no full-body consumption was observed,
so the zero delta does not establish a Skill benefit. The full repository test suite was not run locally; GitHub CI status is reported by this PR's checks.

AI usage statement

OpenAI Codex (GPT-6) assisted implementation and review. Claude Code model invocations supplied the explicitly identified evaluation data.

@Teingi Teingi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two P2 gaps can let the qualification gate pass invalid behavior or incomplete evidence. The attached Run4 sessions themselves are complete; the inline comments describe counterexamples reproduced on separate copies.

Comment on lines +35 to +39
- tool_called_in_turn:
turn: 1
name: mcp__powercontext__remember_memory
args:
scope_id: skill-up-fixture-scope

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject additional writes outside the resolved Scope

In skill-up v0.12.0, this rule succeeds if any call matches the expected arguments. An agent can call remember_memory with scope_id: other-project-scope, then call it with skill-up-fixture-scope, and still pass all the rules despite the instruction to use the resolved Scope for every operation. The reporter also discards arguments: adding the extra wrong-Scope call to a copy of Run4 still produced PASS. Please check the Scope of every relevant recorded operation, so one correctly scoped call cannot hide an unauthorized write.

Comment on lines +137 to +140
if event.get("type") == "assistant":
assistant_turns.add(current_turn)
blocks = content if isinstance(content, list) else []
blocks = [block for block in blocks if isinstance(block, dict) and block.get("type") == "tool_use"]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Require completed turns before marking session evidence complete

Any assistant event satisfies this check, including an unanswered tool call. In a copy of the attached Run4 archive, truncating the explicit-save/with_skill JSONL after line 19 leaves it ending at remember_memory with stop_reason: "tool_use", without the tool result or final answer. Keeping result.json unchanged, report.py still exits successfully with PASS, evidence_complete: true, and no errors. Please require completion evidence for each logical turn, including responses to recorded tool calls, so a truncated archive cannot qualify as complete.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evaluation): evaluate the integration Skill guidance with skill-up

2 participants