Repository navigation
feat(llm): request-side thinking control, effective-policy records, dropped-block diagnostic (#625) - #660
Conversation
…ropped-block diagnostic (#625) Expose the SDK's thinking parameter without changing the default instrument: a provider entry may set "thinking" (verbatim dict — the SDK validates the shape; newer SDKs accept adaptive, older enabled/disabled with a budget), threaded capability-conditionally through build_adapter exactly like #604's request_timeout, with the same one-time-per-type warning when a provider type does not consume it. anthropic + bedrock adapters declare the kwarg; absent (the default) sends NO thinking key — the request stays byte-identical to the pre-change shape, so verdict behavior is unchanged (#242's instrument rule). The effective policy is now visible where decisions read it: - the checkpoint KEY folds a non-None thinking policy into fingerprint_for_binding (a resumed run under a different policy must not adopt stale checkpoints); None folds nothing, so the digest — and existing default-config checkpoints — survive unchanged; - every LLM phase's step report carries {provider, model, thinking} via binding_policy_summary (scanner's seven LLM steps + the CLI's generate-context). The dropped-block diagnostic: the block translation counted, per kind, the response content blocks it drops (thinking, redacted_thinking, refusal...) while usage forwards — CompletionResult.dropped_block_kinds (a COUNT, never a token split; usage cannot split thinking), so content-vs-usage reconciliation has a number to reconcile against. The one-time-per-kind stderr warning stays.
…drops thinking blocks (#625 T1 retro) The retro bug-hunt round found a reachable bug in the thinking control: a thinking-enabled request WITH tools requires the thinking blocks preserved on the echoed assistant turn (Anthropic's documented contract), and this adapter's multi-turn loop echo filters to text/tool-use — so iteration 2 would 400 after iteration 1 is billed. The per-provider knob means one enabled entry hits every tool-using phase (enhance, verify). Refuse the combination loudly before the paid call (both adapters); disabled thinking with tools stays allowed; thinking without tools stays allowed. Three regression tests pin the guard.
|
Collision found and a bug fixed — the retro process review (2026-09-21) surfaced both: 1. This PR collides with #658 (opened 2026-09-20, branch
Recommendation: #658 survives; this PR is superseded. The two pieces worth absorbing into #658 from here: the 2. The retro bug-hunt found a reachable bug in this branch — fixed in the latest push. A thinking-enabled request WITH tools requires thinking blocks preserved on the echoed assistant turn; the multi-turn loop echo filters to text/tool-use, so iteration 2 would 400 after iteration 1 is billed. The per-provider knob meant one enabled entry hit every tool-using phase. The latest commit refuses the combination loudly at build time (both adapters), with three regression tests. If #658 survives instead, its existing guard covers this — but the finding stands for any future reimplementation. |
…ters (#625 review) The committed guard coverage exercised the Anthropic adapter only; the review verified the guard fires in both by manual probe but the suite should pin it — now parameterized over anthropic + bedrock.
|
Correction to an earlier claim in this body: the "one pre-existing failure ( |
What
Exposes request-side thinking control for the Anthropic-format providers and makes the effective policy visible to checkpoint decisions and step reports — plus a dropped-block diagnostic for content-vs-usage reconciliation. Closes #625.
The instrument rule (default unchanged)
A provider entry may now set
thinking:The value passes verbatim as the request's
thinkingparameter (the SDK validates the shape — newer SDKs accept adaptive; older enabled/disabled with a budget). Absent (the default) sends nothinkingkey at all — the request is byte-identical to before, so verdict behavior is unchanged. Threading is capability-conditional exactly like #604'srequest_timeout, with the same one-time-per-type warning when a provider type doesn't consume the knob.Where the effective policy is recorded
fingerprint_for_bindingfolds a non-None policy into the KEY — a resumed run under a different thinking policy will not adopt stale checkpoints (it's a different instrument). ANonepolicy folds nothing, so the digest — and every existing default-config checkpoint — is unchanged by this upgrade.{step}.report.jsoninputs now carry{"provider", "model", "thinking"}(analyze, enhance, verify, report, dynamic-test, app-context, llm-reachability, generate-context), via a singlebinding_policy_summaryhelper read off the live adapter — the record states what was actually in effect.The dropped-block diagnostic
The response translation drops block kinds it can't deliver (thinking, redacted_thinking, refusal) while forwarding their usage. It now counts them per kind on
CompletionResult.dropped_block_kinds— a count, never a token split, since usage cannot split thinking. This gives the reconciliation the issue asked for: delivered content vs N dropped blocks alongside the forwarded usage. The existing one-time-per-kind stderr warning stays.Testing
New regression test (
test_issue625_thinking_control.py, 13 cases): config parse (verbatim + absent-is-None), registry threading + the unconsumed-knob warning, request shape for both adapters (configured key present; default key absent), the dropped-block counts, the fingerprint fold (configured changes the digest; default digest unchanged), and the policy summary. Full adapter/config/identity suites green: 185 passed. The full test directory runs 1974 passed, 148 skipped — one pre-existing failure (test_issue520_blackout_silence.py::test_fold_advisory_reaches_the_envelope_channel) fails identically on pristine master (verified on the clean worktree) and is unrelated to this change.Deferred (on record)
CompletionResultfield is generic and they adopt counting when they carry droppable kinds.Coordination