fix(runtime-host): ask auxiliary calls for the least reasoning each model accepts - #5730
Conversation
…odel accepts Daily Review and prompt suggestions ask for no reasoning with thinkingLevel 'off'. On models that do not declare 'off', resolveThinkingLevel drops it and the request carries no effort at all, so the provider runs its default reasoning (medium on GPT-5, o3 and GPT-6; dynamic on Gemini 3). Prompt suggestions refused every such model instead. leastReasoningThinkingLevel picks 'off' where the model declares it, or where an omitted parameter already means no extended reasoning (no declared levels, Anthropic Messages), and otherwise the model's lowest declared level, which the wire builders send explicitly. Auxiliary calls opt in with reasoning: 'least'; Daily Review and prompt suggestions do. Suggestions no longer refuse reasoning-only models and give them a 1,024-token budget, since reasoning tokens count against it. Fixes apache#5690 Generated-by: Claude Code Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hqhq1025
left a comment
There was a problem hiding this comment.
Reviewed exact head ccac57decd34967ac0deafa8c38c80960eb63008 against base cd94498521df191578a06692f8f87f910673c901. The change adds a shared least-reasoning selector and applies it to Daily Review and prompt-suggestion auxiliary calls, including a larger suggestion output budget when reasoning cannot be disabled. I found one P1 in the Kimi Coding Plan Anthropic route; see the inline comment.\n\nValidation: clean install and build:test; Runtime 3518 pass / 13 skip; Runtime Host 2119 pass / 19 skip with one local sandbox-boundary failure reproduced on the exact base; full typecheck, lint, format, ASF headers, and git diff --check passed. A clean synthetic merge with current main 8d5a3cac3504c36cb1849b060153b58188e2c7b2 also passed build, Runtime, typecheck, lint, format, and ASF checks, with the same base-reproducible Runtime Host sandbox failure. Hosted test and label are green. I did not call the real Kimi service.\n\n> Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
| ): ThinkingLevel { | ||
| const variants = thinkingVariantsForConnection(connection, modelId); | ||
| if (variants.length === 0 || variants.includes('off')) return 'off'; | ||
| if (runtime.wire === 'anthropic-messages') return 'off'; |
There was a problem hiding this comment.
[P1] Do not treat every Anthropic Messages route as reasoning-off capable. kimi-coding-plan uses this wire for k3/k3-256k, but their declared variants are low, high, and max, and the provider-specific builder below defaults an omitted level to max. This branch returns off; the auxiliary call then emits neither thinking nor effort. Through createHostPromptSuggestionModel on this exact head I observed a real K3 request with max_tokens: 128 and no reasoning fields, so the newly enabled paid suggestion path runs at the provider default maximum reasoning under the non-reasoning budget. The exact base rejected the same model before transport creation (fetches = 0). Please choose the lowest declared level for Kimis required-thinking Anthropic route (and other such routes) and add a production-path Kimi regression.
There was a problem hiding this comment.
Confirmed, thanks. The Anthropic Messages exception was too broad. It holds for Claude, where an omitted thinking/effort means no extended thinking. It does not hold for Kimi Coding Plan: k3, k3-256k and kimi-for-coding declare low/high/max and think by default when nothing is sent. Reproduced: on the old head, leastReasoningThinkingLevel returned off for K3 and the options were {}.
Fixed in 00f53a5: the exception is now limited to Claude models on Anthropic Messages (claudeFamilyId(modelId).startsWith('claude-')). Every other route gets its lowest declared level, sent explicitly. For Kimi that is low, with the builder's thinking mode (adaptive for K3, enabled with its 1,024 budget for kimi-for-coding).
Tests:
- Unit:
k3,k3-256kandkimi-for-codingresolve tolow, and the exact Anthropic options are asserted. Claude on bothanthropicandopencodestill resolves tooffwith nothinking/effort. - Production path:
createHostPromptSuggestionModelwith akimi-coding-planconnection onk3. It asserts the request body carriesthinking: { type: 'adaptive' },effort: 'low'andmax_tokens: 1024.
I did not call the real Kimi service either.
…ning The Anthropic Messages exception in leastReasoningThinkingLevel held for Claude, where an omitted thinking/effort means no extended thinking, but not for Kimi Coding Plan: k3, k3-256k and kimi-for-coding think by default when no level is sent, so a least-reasoning call ran at the route's default thinking. Keep 'off' only for Claude on Anthropic Messages; every other route gets its lowest declared level explicitly. Generated-by: Claude Code Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hqhq1025
left a comment
There was a problem hiding this comment.
Re-review result
Exact head: 00f53a59b5a3f699913cd4e4c2b8f37b234891dc
I found no P0-P3 issue on this head. The previous Kimi Coding Plan P1 is fixed.
The new commit narrows the Anthropic Messages no-reasoning exception to Claude model families. Kimi Coding Plan k3, k3-256k, and kimi-for-coding now resolve the auxiliary-call request to their lowest declared effort, low, instead of omitting the setting and falling back to the route's maximum reasoning (packages/runtime/src/model-factory.ts:455-486). The unit regression checks all three Kimi model shapes while preserving omitted reasoning for Claude on native and gateway Anthropic routes (packages/runtime/src/__tests__/model-factory-thinking.test.ts:1333-1372). The production Host regression confirms a K3 prompt-suggestion request carries adaptive thinking, effort: low, and the reasoning-sized output budget (packages/runtime-host/src/__tests__/execution-model-composition.test.ts:6334-6432).
I also exercised kimi-for-coding through the same production Host path: it emitted enabled thinking plus effort: low; the Anthropic SDK expanded the wire max_tokens to 2,048 to accommodate the 1,024-token thinking budget. The new K3 production regression fails on the prior head ccac57de because the request has no thinking, and passes on this head.
Validation:
- Clean
npm ciandnpm run build:testwith Node 24.18.1. - Runtime: 3,519 passed / 13 skipped.
- Runtime Host: 2,120 passed / 19 skipped / 1 failed. The sole managed-sandbox boundary failure is the same environment failure previously reproduced on the exact base and is outside the changed path.
- Focused model-factory suite: 53/53; focused production Host suggestion cases: 2/2.
- Full typecheck, lint, format, ASF headers, model-metadata check, protocol epoch guard, and
git diff --checkpassed. - Hosted
testis green. - Current main is
87fc9f69cd11648f31048c1b633bb813aca5f515. The PR is 2 commits ahead and 2 behind; merge-tree is conflict-free. Synthetic merge45fd9b1c748696077024071e3992f7edb6dce822passed clean build, focused regressions, typecheck, lint, format, ASF headers, protocol epoch guard, and diff check.
Residual scope: I did not call the real Kimi service or run packaged/native Windows or macOS validation.
Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
likun666661
left a comment
There was a problem hiding this comment.
Approved at head 00f53a5. The resolved-target least-reasoning selection is passed through to provider options, the suggestion budget accounts for required reasoning, and the Kimi Anthropic-route regression is covered at both the selector and Host request levels. I found no blocking issue in this scoped change.\n\nIntegration note (not a code blocker here): #5723 is still open, so Codex subscription auxiliary calls need that fix before this behavior works end to end on that connection. The existing five-second suggestion deadline can also discard slower responses; that is a separate product follow-up, not something this PR claims to solve. I inspected the current CLI diff and green hosted CI, but did not rerun the suites locally.
Watermark bfb315a. Done: apache#5573/apache#5600/apache#5601, apache#5521, apache#4875, apache#5723, apache#5738, apache#5742. Not applicable: apache#5737, apache#5593. Deferred: apache#5730. Consider: apache#5599, apache#5120, apache#5693. Diverged: apache#5740. Skipped: ACP, WorkHub, upstream renderer and packages/ui, one refactor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Summary
Daily Review and next-prompt suggestions ask for no reasoning by setting
thinkingLevel: 'off'. On models that do not declareoff,resolveThinkingLeveldrops it, and the request then carries no effort at all. The provider falls back to its default reasoning: medium on GPT-5, o3 and GPT-6, and dynamic on Gemini 3.#5629 worked around this for suggestions by refusing every such model, so on those models suggestions never appear. That includes GPT-6 Sol, Luna and Astra, GPT-5, o3/o4-mini and Gemini 3.x.
This PR adds
leastReasoningThinkingLevelnext to the wire builders inmodel-factory.ts. It picks the level that asks a model for the least reasoning it accepts:'off'when the model declares it;'off'when the model declares no levels, or speaks Anthropic Messages, because there an omitted parameter already means no extended reasoning;minimal/low), which the wire builders then send explicitly.runHostAuxiliaryModelCallgainsreasoning: 'least', which resolves this against the target model. Daily Review and prompt suggestions use it.Suggestions no longer refuse reasoning-only models. When a model can't turn reasoning off, their output budget grows from 128 to 1,024 tokens, because reasoning tokens count against it.
Behaviour changes:
Fixes #5690
Verification
leastReasoningThinkingLevelis covered per provider, including the wire field it sends:gpt-5→minimal;o3,gpt-6-soland Codexgpt-6-astra→low;gemini-3.5-flash→minimal;gpt-5.5staysnoneandgemini-2.5-flashstays budget 0;off, sending noeffortorthinking.gpt-5sendsreasoning.effort: 'minimal'withmax_output_tokens: 1024, whileoff-capable models still sendnonewith 128.gpt-6-astra, three turns, thensession.prompt-suggestion.generateafter each:main:nonein about 0.4 s, and no request sent. The model is refused before the call.success, 313 in / 11 out, 0 reasoning tokens atlow, 3.7 s).runtimeandruntime-host:tsc --noEmitpasses,biome checkpasses on the changed files, and so do the touched suites (model-factory-thinking,prompt-suggestion, and the affectedexecution-model-compositioncases).rmto tolerateEBUSYfrom SQLite handles. That is a local test-harness issue and is not changed here.biome.cmdspawnSync→EINVAL), so I ran its steps manually.Review focus
gpt-6-astraatlow, with 0–17 reasoning tokens.AI use
Tool(s) and scope: Claude Code (Claude Opus) implemented the change and tests and ran the verification above; liugddx reviewed. The commit carries a
Generated-bytrailer.Checklist
Does this PR entail a change in behavior?
🤖 Generated with Claude Code