Version v3.86.0
Problem (one or two sentences)
Zoo Code splits a command into sub-commands before matching them against the allow/deny lists, and that splitter is built from regular expressions that cannot see nesting. A command substitution hidden inside a parameter expansion (echo ${x:-$(rm -rf ~)}) is therefore never decomposed, and a nested substitution ($(foo $(bar))) is truncated at the first ) — in both shapes the inner payload is never matched against the lists, and the shell runs it on the outer command's approval.
Context (who is affected and when)
Every user who relies on the allowlist for hands-free operation. The expectation that inner commands are evaluated on their own is pinned by an existing test: echo $(whoami) must return ask_user when whoami is not on the allowlist (commands.spec.ts:94-96). That expectation is the independent-matching rule — a sub-shell payload needs its own allowlist match to be approved, and it is pinned only by that test, while the parser-level split it depends on is separately pinned (parse-command.spec.ts:181-186: echo $(whoami) yields a distinct whoami segment). The two bypasses below break the rule for shapes the regex pipeline mis-masks: anything an agent — or a prompt-injected model — can emit that the parser cannot fully account for is approved or denied on a false decomposition.
Reproduction steps
The two snippets below are paste-ready for commands.spec.ts.
it("escapes substitution extraction: expansion mask swallows the nested $(...)", () => {
// parseCommand returns ONE segment — the whole string — so the payload
// is never matched against the lists. `echo` is allowlisted, but the
// shell still runs `rm -rf ~/work`, because ${x:-...} falls back to the
// substitution when x is unset. Observed today: auto_approve. Expected
// per the independent-matching rule pinned at commands.spec.ts:94-96:
// ask_user.
expect(getCommandDecision("echo ${x:-$(rm -rf ~/work)}", ["echo"])).not.toBe("auto_approve")
})
it("truncates on nesting: the payload rides two outer allowlist entries", () => {
// parseCommand returns ["ls", "foo $(rm -rf /", "bar )"] for
// `ls $(foo $(rm -rf /) bar)`: the lazy $\((.*?)\) regex cuts at the
// first ) and `rm -rf /` never becomes its own segment. The truncated
// middle rides the `foo` allowlist entry, the trailing fragment rides
// `bar`. Observed today: auto_approve.
expect(getCommandDecision("ls $(foo $(rm -rf /) bar)", ["ls", "foo", "bar"])).not.toBe("auto_approve")
})
Both assertions fail on current main: each call returns auto_approve, with the segment lists shown in the comments.
Root cause: mask-then-split, in the wrong order, with non-nesting regexes
The parseCommand docstring says the parser handles subshell forms ($(cmd), `cmd`, <(cmd), >(cmd)), but parseCommandLine pre-replaces shell constructs with placeholder tokens (__PARAM_0__, __SUBSH_0__, …) before handing the string to shell-quote:
- Mask order hides substitutions. Parameter expansions
\$\{[^}]+\} are masked before command-substitution extraction (parse-command.ts:575) — the \$\((.*?)\) pass only runs later (:622). The $(rm -rf ~/work) inside ${x:-$(rm -rf ~/work)} is already behind a placeholder when extraction runs, so it is restored verbatim into one echo-prefixed segment and never matched against the lists.
- No paren balancing. The substitution regex is lazy —
\$\((.*?)\) truncates $(foo $(bar)) at the first ), extracting foo $(bar as the sub-shell text; the inner command never becomes its own evaluated segment. Process substitutions [<>]\(([^)]+)\) (:581) cannot nest either.
- The blacklist is a race it cannot win.
containsDangerousSubstitution enumerates known bash/zsh expansion gadgets by regex — ${var@P}, octal-escape assignments, zsh glob qualifiers — and a construct outside the catalogue passes: echo ${a:-$(evil)} returns auto_approve with an echo-only allowlist. A denylist of syntax cannot substitute for a parser that accounts for all syntax.
The pipeline is quote-aware — the shared state-machine scanner scanTopLevelQuotes tracks quotes, heredocs, and comments, and the split uses its output via maskTopLevelQuotes — but it does not track parenthesis or brace depth, so nesting is as invisible to it as to the regexes.
Expected result
- Every command substitution — nested, or embedded in a parameter expansion, process substitution, or heredoc — is extracted with balanced-paren awareness, and each inner command independently matches the lists (the independent-matching rule pinned for the simple case at
commands.spec.ts:94-96).
- Fail closed: if the parser cannot fully account for the input's structure — unbalanced delimiters, unclosed expansions — the result is
malformed_command (ask/deny), never an auto_approve built on a partial decomposition. Today the fail-closed path fires only for unterminated quotes and heredocs (parse-command.ts:479-487, commands.ts:314-316).
Hardening direction
Replace the sequential regex pre-replacement pipeline with one quote-aware, left-to-right scan that simultaneously tracks quote state, parenthesis/brace depth, and here-doc state, emitting sub-commands and placeholders from a single pass. scanTopLevelQuotes is the right starting point — it already handles quotes, heredocs, and comments correctly — but it would need paren/brace-depth tracking added. The fail-closed rule above makes parser gaps safety-neutral while the rewrite lands: an unparseable-but-plausible command must cost a prompt, not a free pass.
Prior discussion
- Issue #1569 / PR #1760 (blanket auto-deny) deliberately carved this out — the PR description's scope note: "hardening the allowlist parser itself (shell expansions and chaining forms the current matcher does not flag) is not part of this PR — follow-up territory."
- The same review thread probed whether the dangerous-substitution blacklist runs twice in the new blanket-deny path (
discussion_r4100696287). A parser that accounts for all syntax would make that blacklist layer redundant — and the re-check unnecessary.
- Related but distinct: whether an allowlist entry like
ls should approve ls -la (whole-command vs. prefix matching) is a matching-semantics question this issue deliberately leaves open; this issue is only about where the string is cut.
Why this is worth fixing
In a hands-free session the allowlist is the user's only defense, and its core promise — every segment matches, or you get asked — is enforced by a regex pipeline that loses segments it cannot see. Both bypass commands above are one-liners a malicious instruction in a repo file or web page could plant in an agent's context, and with execute auto-approval plus the Destructive Command Guard enabled and returning an allow verdict, the whole allow/deny pipeline — blacklist included — is bypassed (index.ts:270-304), so nothing behind the parser catches these shapes either. The fix is bounded: one file plus a fail-closed rule in the decision aggregation — and every bypass in this report is a ready-made regression test.
Version v3.86.0
Problem (one or two sentences)
Zoo Code splits a command into sub-commands before matching them against the allow/deny lists, and that splitter is built from regular expressions that cannot see nesting. A command substitution hidden inside a parameter expansion (
echo ${x:-$(rm -rf ~)}) is therefore never decomposed, and a nested substitution ($(foo $(bar))) is truncated at the first)— in both shapes the inner payload is never matched against the lists, and the shell runs it on the outer command's approval.Context (who is affected and when)
Every user who relies on the allowlist for hands-free operation. The expectation that inner commands are evaluated on their own is pinned by an existing test:
echo $(whoami)must returnask_userwhenwhoamiis not on the allowlist (commands.spec.ts:94-96). That expectation is the independent-matching rule — a sub-shell payload needs its own allowlist match to be approved, and it is pinned only by that test, while the parser-level split it depends on is separately pinned (parse-command.spec.ts:181-186:echo $(whoami)yields a distinctwhoamisegment). The two bypasses below break the rule for shapes the regex pipeline mis-masks: anything an agent — or a prompt-injected model — can emit that the parser cannot fully account for is approved or denied on a false decomposition.Reproduction steps
The two snippets below are paste-ready for
commands.spec.ts.Both assertions fail on current
main: each call returnsauto_approve, with the segment lists shown in the comments.Root cause: mask-then-split, in the wrong order, with non-nesting regexes
The
parseCommanddocstring says the parser handles subshell forms ($(cmd),`cmd`,<(cmd),>(cmd)), butparseCommandLinepre-replaces shell constructs with placeholder tokens (__PARAM_0__,__SUBSH_0__, …) before handing the string toshell-quote:\$\{[^}]+\}are masked before command-substitution extraction (parse-command.ts:575) — the\$\((.*?)\)pass only runs later (:622). The$(rm -rf ~/work)inside${x:-$(rm -rf ~/work)}is already behind a placeholder when extraction runs, so it is restored verbatim into oneecho-prefixed segment and never matched against the lists.\$\((.*?)\)truncates$(foo $(bar))at the first), extractingfoo $(baras the sub-shell text; the inner command never becomes its own evaluated segment. Process substitutions[<>]\(([^)]+)\)(:581) cannot nest either.containsDangerousSubstitutionenumerates known bash/zsh expansion gadgets by regex —${var@P}, octal-escape assignments, zsh glob qualifiers — and a construct outside the catalogue passes:echo ${a:-$(evil)}returnsauto_approvewith anecho-only allowlist. A denylist of syntax cannot substitute for a parser that accounts for all syntax.The pipeline is quote-aware — the shared state-machine scanner
scanTopLevelQuotestracks quotes, heredocs, and comments, and the split uses its output viamaskTopLevelQuotes— but it does not track parenthesis or brace depth, so nesting is as invisible to it as to the regexes.Expected result
commands.spec.ts:94-96).malformed_command(ask/deny), never anauto_approvebuilt on a partial decomposition. Today the fail-closed path fires only for unterminated quotes and heredocs (parse-command.ts:479-487,commands.ts:314-316).Hardening direction
Replace the sequential regex pre-replacement pipeline with one quote-aware, left-to-right scan that simultaneously tracks quote state, parenthesis/brace depth, and here-doc state, emitting sub-commands and placeholders from a single pass.
scanTopLevelQuotesis the right starting point — it already handles quotes, heredocs, and comments correctly — but it would need paren/brace-depth tracking added. The fail-closed rule above makes parser gaps safety-neutral while the rewrite lands: an unparseable-but-plausible command must cost a prompt, not a free pass.Prior discussion
discussion_r4100696287). A parser that accounts for all syntax would make that blacklist layer redundant — and the re-check unnecessary.lsshould approvels -la(whole-command vs. prefix matching) is a matching-semantics question this issue deliberately leaves open; this issue is only about where the string is cut.Why this is worth fixing
In a hands-free session the allowlist is the user's only defense, and its core promise — every segment matches, or you get asked — is enforced by a regex pipeline that loses segments it cannot see. Both bypass commands above are one-liners a malicious instruction in a repo file or web page could plant in an agent's context, and with execute auto-approval plus the Destructive Command Guard enabled and returning an allow verdict, the whole allow/deny pipeline — blacklist included — is bypassed (
index.ts:270-304), so nothing behind the parser catches these shapes either. The fix is bounded: one file plus a fail-closed rule in the decision aggregation — and every bypass in this report is a ready-made regression test.