Worktree a2 h3 - #4
Merged
Merged
Conversation
Track A2 object H3 proposes bounding old-coder-adversary's Bash by a git-only grammar. Step zero measured what the reviewer actually runs: 57 Bash calls across 8 recorded sessions, selected from subagent metadata rather than recalled. The sketch admits 8 of them. The SPEC argues a wider grammar from those rows and carries four rulings for the human. The sharpest: `;` appears in 22 of 57 commands as read sequencing, so the object's own acceptance criterion naming it an excluded class cannot hold alongside "the grammar covers observed use". The harvester ships beside its output so a reader on another host can measure their own reviewer instead of trusting these eight sessions. Home paths are rewritten, both the path and the slug spelling, and the redaction is declared; verified not to change any grammar verdict. AWAITING APPROVAL. No mechanism yet, no audit row moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All four answered as proposed: chains allowed with each segment checked, Python denied with no carve-out, the two stderr redirect spellings approved, and the three audit rows may move on a write-capability-only claim. Decide 2 keeps the table that answered it. Seven of the reviewer's eight Python calls were Python used as a richer grep and the grammar reaches all seven. The eighth re-ran a repo check to test a claimed-green row, and nothing in the grammar reaches it. The carve-out that would have was refused: it would run a file out of the tree under review, which is the hazard EX-8 already names. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
110 cases written from docs/spec-a2-h3.md before any handler exists: 14 in grammar, 30 across the four excluded classes, 6 fail-closed paths, one that proves the handler declines to decide about tools other than Bash, and the 57-row harvest as the positive control. Watched failing against a handler that exits 0 and says nothing: 51 of 110 fail. The 59 that pass are the allow cases, which a permissive stub also allows, so the suite discriminates rather than merely running. expected.tsv labels every harvested row and is authored from the SPEC, not read back out of the handler. Its thirteen denials name the class each falls in so a reader can check them one by one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
hooks/adversary-bash-grammar.py bounds old-coder-adversary's Bash to an allowlist over reads. It never asks whether a string writes, which is undecidable; it asks whether a string is one of a small number of written-down shapes and refuses everything else, including every string whose effect it cannot determine. 115 controls, and the layer runs them twice: once against the handler and once against a stub that allows everything, where the suite must fail. The first wall. tools/audit_sweep.py said a hook never lifts a shell tool, and said a future object bounding a shell by an allowlist grammar would change that deliberately and say so. This is it saying so. Four conditions now, all required: an exact matcher, a handler this module names in ALLOWLIST_SHELL_HANDLERS, a probe record carrying the handler's current sha256, and that record declaring `grammar: allowlist`. Six new controls, one per way of earning the lift without meeting it. The docstring states what the constant proves and what it does not: it stops the accident, not the contributor who means to mislead. Two limits found while building and written down rather than left for the round. git honours diff.external and .gitattributes textconv from the repository under review, so `git diff` can run code the grammar never sees; closing that needs a rewriting hook, which clause 1 of the tier's test puts outside hooks/. And ruff and mypy never reach hooks/ or tools/, so the handler ships held by its behavioural cases alone. Not yet registered in the adversary's frontmatter. The round that grades this object reviews the bound on itself, so it runs first, unbounded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adversarial round bound to 4c34bde...19a2a36, tree f4543ac8af211553, run before the hook was registered so the reviewer grading this object still held unbounded Bash. Two findings, no escape from the grammar. Upheld: tail -F was allowed. The grammar refuses tail -f because it never terminates, and -F is GNU's --follow=name --retry. Fixed as a class rather than a spelling, since short options cluster and -fn10 follows too. Rejected with evidence: find -print0 and find -- are denied and should be. Neither appears in the 57-row harvest, and the round's own example pipeline is denied again at xargs, so admitting -print0 would not make it work. The grammar covers observed use; widening it for a command nobody ran is how an allowlist decays into a list of things somebody thought of. Closed an observation the round raised but did not count: audit_sweep resolved handlers by stem, so a frontmatter naming .sh when only .py exists would credit a lift for a file the runtime cannot run. hooks_registered already reddens on that, but this object's own suffix probing created the disagreement, so handler_path now honours an explicit suffix. bash-grammar-controls 118, audit-sweep-controls 24. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The adversarial round's coverage section listed what it had not reached. Two of those checks, run afterwards, found defects the findings section did not. The handler was wider than docs/spec-a2-h3.md. FIND_PREDICATES carried -mindepth, -ipath, -prune and -print; PLAIN_READERS carried md5sum, basename, dirname, pwd and date. None is in the approved table, none is in the harvest, and none had a control, so the widening was neither ruled on nor measured nor tested. Each is harmless, which is how they got there, and which is the argument this object rejected an hour earlier when it refused find -print0. Narrowed to the approved table with four controls pinning the edge. expected.tsv labelled four rows as shell-variable denials. The rule that fires is the multiline check. The denials were over-determined so no case failed and nothing caught it, which is a control agreeing with the code by accident. The labels now name the operative rule. bash-grammar-controls 122. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two rulings from the human after the adversarial round. -print and -print0 join the find table. Both change the output separator and reach nothing. They were refused when the round raised them, because the rule then was that the grammar covers observed use and neither is in the harvest. They are in now because somebody ruled on them, which is the only thing that moves the table. -prune and -mindepth stayed out, with controls proving the amendment widened it by exactly two entries rather than opening it. agents/old-coder-adversary.md gains its frontmatter hooks block, after the round, so the reviewer that graded this object held unbounded Bash throughout. hooks-registered now sees two handlers and reports the new one NOT installed on this host, a note rather than a failure: the tier is opt-in and the install is a documented step. bash-grammar-controls 124. Still pending the host probe, which is the only thing that proves the runtime calls the handler. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Host probe 1 came back void on both controls and is not recorded. The install cause is development-time; the probe-design cause would have produced a green record for an unproven hook had the install been right. Both are written up as Decide 5 and Decide 6 rather than left in chat.
The probe's negative control moves from a write to a redirect, because sed -i is refused by the reviewer's own brief before the grammar is ever consulted. Step 2 now requires the handler's own denial text and treats a bare refusal as void rather than a pass. hooks_registered gains installed_agent_note: the handler resolves by a host-absolute address, the agent file does not, and a <config>/agents symlink aimed at another checkout makes every check green while the subagent spawns unbounded. Two controls, non-vacuous against a stub. The handler is untouched, so its sha256 is unchanged.
The handler's atime predated the probe and /tmp/h3-probe.txt was never written, so neither control reached it. Both used git, which rtk and the worktree guard refuse from a worktree-isolated session ahead of any agent-level hook. Carrier commands move to rg and sed -n, which nothing upstream objects to. The class stays a redirect, so this is inside ruling 5. The record now carries the handler's atime, which answers the one question hooks_registered says it cannot.
Retracts the atime evidence from probe 2: /home/mcrowe/Programming is mounted noatime, so access times never update and the observation was consistent with both outcomes. The mount was read after the claim. Probe 3's negative control was refused by the reviewer's own prompt defense, the same signature as probe 1. That is twice, so the object stops rather than earning a fourth attempt. The reason is structural: the grammar was built to never deny normal review work, so no command the reviewer issues on its own is refused, and any command delivered through its brief is refused as injection before the handler is consulted.
The stable-failure rule stops a mechanism that fails the same way twice. Probes 1 and 3 never reached the mechanism: both refused a write-looking command delivered as a peer message. That is a probe-design failure with an untried alternative, so the stop was wrong. Session hooks are cleared as a cause by the reference and by running them on the probe payloads: parallel, original input each, no decision from rtk or context-mode. The likely cause of probes 2 and 3 is folder trust: since 2.1.218 a frontmatter hook is silently skipped in an untrusted folder, a parent's trust does not carry down, and the worktree has no trust entry. Probe 4: pwd elicited as a question, spawned directly, sed -n positive control, session under --debug, worktree trusted first. Two rulings opened: a trust check in hooks_registered, and the composition limit.
Probe 4 with the worktree trusted and --debug on shows the Agent tool spawns old-coder-adversary as an in-process teammate. Its three Bash calls fired no PreToolUse hook at all, not the grammar and not rtk. The in-process runner logs Found 0 total hooks in registry. The frontmatter hook is never invoked for the adversary on this runtime, so the bound cannot be proven here. Stable failure for want of runtime scope: VE-1, EX-5, EX-7 stay delegated to the runtime repo with a proven reason. The hook and its 124 CI controls stay landed as the portable artifact.
Status table records the outcome: grammar handler, both walls, and 160 controls landed and green, one adversarial round, but four host probes proved a frontmatter hook cannot bound an in-process teammate on this runtime. VE-1, EX-5, EX-7 stay delegated. Hook and controls remain as the portable artifact.
H3's hooks-tier tooling moved the demo tree to ffda7020f2c961fc. Rebind the source-state pair (ec89de1 / ffda7020) so evidence-binding is green, and reframe the toolchain paragraph honestly: CI verified the earlier tree 0f6d9e8 at ad2b4a0; the current tree is local-gauntlet green on 3.14.7 with no CI run yet, which the declared downgrade already covers. All 22 layers green.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.