Skip to content

fix(runtime-host): ask auxiliary calls for the least reasoning each model accepts - #5730

Merged
liugddx merged 2 commits into
apache:mainfrom
liugddx:fix/aux-minimal-reasoning
Sep 26, 2026
Merged

liugddx merged 2 commits into
apache:mainfrom
liugddx:fix/aux-minimal-reasoning

Conversation

@liugddx

@liugddx liugddx commented Sep 25, 2026

Copy link
Copy Markdown
Member

Summary

Daily Review and next-prompt suggestions ask for no reasoning by setting thinkingLevel: 'off'. On models that do not declare off, resolveThinkingLevel drops it, and the request then carries no effort at all. The provider falls back to its default reasoning: medium on GPT-5, o3 and GPT-6, and dynamic on Gemini 3.

#5629 worked around this for suggestions by refusing every such model, so on those models suggestions never appear. That includes GPT-6 Sol, Luna and Astra, GPT-5, o3/o4-mini and Gemini 3.x.

This PR adds leastReasoningThinkingLevel next to the wire builders in model-factory.ts. It picks the level that asks a model for the least reasoning it accepts:

  • 'off' when the model declares it;
  • also 'off' when the model declares no levels, or speaks Anthropic Messages, because there an omitted parameter already means no extended reasoning;
  • otherwise the model's lowest declared level (minimal/low), which the wire builders then send explicitly.

runHostAuxiliaryModelCall gains reasoning: 'least', which resolves this against the target model. Daily Review and prompt suggestions use it.

Suggestions no longer refuse reasoning-only models. When a model can't turn reasoning off, their output budget grows from 128 to 1,024 tokens, because reasoning tokens count against it.

Behaviour changes:

  • Daily Review on those models now sends an explicit lowest effort instead of running at the provider default.
  • Suggestions are now offered on them. That adds a paid call per reply for users who opted in.

Fixes #5690

Verification

  • leastReasoningThinkingLevel is covered per provider, including the wire field it sends:
    • gpt-5 → minimal; o3, gpt-6-sol and Codex gpt-6-astra → low; gemini-3.5-flash → minimal;
    • gpt-5.5 stays none and gemini-2.5-flash stays budget 0;
    • Anthropic effort models and undeclared models keep off, sending no effort or thinking.
  • The composition test now asserts the request body: gpt-5 sends reasoning.effort: 'minimal' with max_output_tokens: 1024, while off-capable models still send none with 128.
  • Real backend. I ran a Runtime Host on a copy of a Desktop workspace with a ChatGPT-subscription connection. This was measured with fix(runtime): stream and fold non-streaming Codex OAuth calls #5723 applied, because without it every Codex auxiliary call fails. Model gpt-6-astra, three turns, then session.prompt-suggestion.generate after each:
    • main: none in about 0.4 s, and no request sent. The model is refused before the call.
    • This branch: a suggestion appears, e.g. "Yes, give me a checklist." (success, 313 in / 11 out, 0 reasoning tokens at low, 3.7 s).
    • With a 30 s deadline, all three calls came back in 3.2–8.3 s with 0–17 reasoning tokens.
  • runtime and runtime-host: tsc --noEmit passes, biome check passes on the changed files, and so do the touched suites (model-factory-thinking, prompt-suggestion, and the affected execution-model-composition cases).
  • On Windows, the composition cases need their temp-root rm to tolerate EBUSY from SQLite handles. That is a local test-harness issue and is not changed here.
  • I did not run the full workspace suites.
  • The pre-commit hook cannot run on Windows (biome.cmd spawnSync → EINVAL), so I ran its steps manually.

Review focus

  • Suggestion deadline. feat(desktop): add opt-in next prompt suggestions #5629's 5 s deadline is shorter than Codex's measured latency. With the default deadline, two of the three real calls above timed out and were discarded, so on Codex some suggestions will still not show. I have left that product value alone. It seems like a separate decision.
  • Output budget. The 1,024-token budget is headroom, not measured need. I only measured gpt-6-astra at low, with 0–17 reasoning tokens.
  • WorkHub routing. Intent and recall still reason at the Coordination Session's own level with an 80-token budget. I did not touch that here.

AI use

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: Claude Code (Claude Opus) implemented the change and tests and ran the verification above; liugddx reviewed. The commit carries a Generated-by trailer.

Checklist

  • Tests cover the change and fail without it
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

🤖 Generated with Claude Code

…odel accepts

Daily Review and prompt suggestions ask for no reasoning with
thinkingLevel 'off'. On models that do not declare 'off',
resolveThinkingLevel drops it and the request carries no effort at all,
so the provider runs its default reasoning (medium on GPT-5, o3 and
GPT-6; dynamic on Gemini 3). Prompt suggestions refused every such
model instead.

leastReasoningThinkingLevel picks 'off' where the model declares it, or
where an omitted parameter already means no extended reasoning (no
declared levels, Anthropic Messages), and otherwise the model's lowest
declared level, which the wire builders send explicitly. Auxiliary calls
opt in with reasoning: 'least'; Daily Review and prompt suggestions do.
Suggestions no longer refuse reasoning-only models and give them a
1,024-token budget, since reasoning tokens count against it.

Fixes apache#5690

Generated-by: Claude Code
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added the effort/M Under 500 readable lines label Sep 25, 2026

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head ccac57decd34967ac0deafa8c38c80960eb63008 against base cd94498521df191578a06692f8f87f910673c901. The change adds a shared least-reasoning selector and applies it to Daily Review and prompt-suggestion auxiliary calls, including a larger suggestion output budget when reasoning cannot be disabled. I found one P1 in the Kimi Coding Plan Anthropic route; see the inline comment.\n\nValidation: clean install and build:test; Runtime 3518 pass / 13 skip; Runtime Host 2119 pass / 19 skip with one local sandbox-boundary failure reproduced on the exact base; full typecheck, lint, format, ASF headers, and git diff --check passed. A clean synthetic merge with current main 8d5a3cac3504c36cb1849b060153b58188e2c7b2 also passed build, Runtime, typecheck, lint, format, and ASF checks, with the same base-reproducible Runtime Host sandbox failure. Hosted test and label are green. I did not call the real Kimi service.\n\n> Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

Comment thread packages/runtime/src/model-factory.ts Outdated
): ThinkingLevel {
const variants = thinkingVariantsForConnection(connection, modelId);
if (variants.length === 0 || variants.includes('off')) return 'off';
if (runtime.wire === 'anthropic-messages') return 'off';

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Do not treat every Anthropic Messages route as reasoning-off capable. kimi-coding-plan uses this wire for k3/k3-256k, but their declared variants are low, high, and max, and the provider-specific builder below defaults an omitted level to max. This branch returns off; the auxiliary call then emits neither thinking nor effort. Through createHostPromptSuggestionModel on this exact head I observed a real K3 request with max_tokens: 128 and no reasoning fields, so the newly enabled paid suggestion path runs at the provider default maximum reasoning under the non-reasoning budget. The exact base rejected the same model before transport creation (fetches = 0). Please choose the lowest declared level for Kimis required-thinking Anthropic route (and other such routes) and add a production-path Kimi regression.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed, thanks. The Anthropic Messages exception was too broad. It holds for Claude, where an omitted thinking/effort means no extended thinking. It does not hold for Kimi Coding Plan: k3, k3-256k and kimi-for-coding declare low/high/max and think by default when nothing is sent. Reproduced: on the old head, leastReasoningThinkingLevel returned off for K3 and the options were {}.

Fixed in 00f53a5: the exception is now limited to Claude models on Anthropic Messages (claudeFamilyId(modelId).startsWith('claude-')). Every other route gets its lowest declared level, sent explicitly. For Kimi that is low, with the builder's thinking mode (adaptive for K3, enabled with its 1,024 budget for kimi-for-coding).

Tests:

  • Unit: k3, k3-256k and kimi-for-coding resolve to low, and the exact Anthropic options are asserted. Claude on both anthropic and opencode still resolves to off with no thinking/effort.
  • Production path: createHostPromptSuggestionModel with a kimi-coding-plan connection on k3. It asserts the request body carries thinking: { type: 'adaptive' }, effort: 'low' and max_tokens: 1024.

I did not call the real Kimi service either.

…ning

The Anthropic Messages exception in leastReasoningThinkingLevel held for
Claude, where an omitted thinking/effort means no extended thinking, but
not for Kimi Coding Plan: k3, k3-256k and kimi-for-coding think by
default when no level is sent, so a least-reasoning call ran at the
route's default thinking. Keep 'off' only for Claude on Anthropic
Messages; every other route gets its lowest declared level explicitly.

Generated-by: Claude Code
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review result

Exact head: 00f53a59b5a3f699913cd4e4c2b8f37b234891dc

I found no P0-P3 issue on this head. The previous Kimi Coding Plan P1 is fixed.

The new commit narrows the Anthropic Messages no-reasoning exception to Claude model families. Kimi Coding Plan k3, k3-256k, and kimi-for-coding now resolve the auxiliary-call request to their lowest declared effort, low, instead of omitting the setting and falling back to the route's maximum reasoning (packages/runtime/src/model-factory.ts:455-486). The unit regression checks all three Kimi model shapes while preserving omitted reasoning for Claude on native and gateway Anthropic routes (packages/runtime/src/__tests__/model-factory-thinking.test.ts:1333-1372). The production Host regression confirms a K3 prompt-suggestion request carries adaptive thinking, effort: low, and the reasoning-sized output budget (packages/runtime-host/src/__tests__/execution-model-composition.test.ts:6334-6432).

I also exercised kimi-for-coding through the same production Host path: it emitted enabled thinking plus effort: low; the Anthropic SDK expanded the wire max_tokens to 2,048 to accommodate the 1,024-token thinking budget. The new K3 production regression fails on the prior head ccac57de because the request has no thinking, and passes on this head.

Validation:

  • Clean npm ci and npm run build:test with Node 24.18.1.
  • Runtime: 3,519 passed / 13 skipped.
  • Runtime Host: 2,120 passed / 19 skipped / 1 failed. The sole managed-sandbox boundary failure is the same environment failure previously reproduced on the exact base and is outside the changed path.
  • Focused model-factory suite: 53/53; focused production Host suggestion cases: 2/2.
  • Full typecheck, lint, format, ASF headers, model-metadata check, protocol epoch guard, and git diff --check passed.
  • Hosted test is green.
  • Current main is 87fc9f69cd11648f31048c1b633bb813aca5f515. The PR is 2 commits ahead and 2 behind; merge-tree is conflict-free. Synthetic merge 45fd9b1c748696077024071e3992f7edb6dce822 passed clean build, focused regressions, typecheck, lint, format, ASF headers, protocol epoch guard, and diff check.

Residual scope: I did not call the real Kimi service or run packaged/native Windows or macOS validation.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

@likun666661 likun666661 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved at head 00f53a5. The resolved-target least-reasoning selection is passed through to provider options, the suggestion budget accounts for required reasoning, and the Kimi Anthropic-route regression is covered at both the selector and Host request levels. I found no blocking issue in this scoped change.\n\nIntegration note (not a code blocker here): #5723 is still open, so Codex subscription auxiliary calls need that fix before this behavior works end to end on that connection. The existing five-second suggestion deadline can also discard slower responses; that is a separate product follow-up, not something this PR claims to solve. I inspected the current CLI diff and green hosted CI, but did not rerun the suites locally.

@liugddx
liugddx merged commit 17d7fe6 into apache:main Sep 26, 2026
1 check passed
Shouly pushed a commit to Shouly/maka that referenced this pull request Sep 27, 2026
Watermark bfb315a. Done: apache#5573/apache#5600/apache#5601, apache#5521, apache#4875, apache#5723, apache#5738,
apache#5742. Not applicable: apache#5737, apache#5593. Deferred: apache#5730. Consider: apache#5599,
apache#5120, apache#5693. Diverged: apache#5740. Skipped: ACP, WorkHub, upstream renderer and
packages/ui, one refactor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/M Under 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(runtime-host): auxiliary calls that ask for thinkingLevel 'off' still run default reasoning on models without an off switch

3 participants