Skip to content

fix: treat llama.cpp exceed-context errors as overflow - #939

Open
v0ropaev wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
v0ropaev:fix/llama-server-context-overflow
Open

v0ropaev wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
v0ropaev:fix/llama-server-context-overflow

Conversation

@v0ropaev

@v0ropaev v0ropaev commented Oct 7, 2026 •

Copy link
Copy Markdown

Closes #937.

llama-server answers an oversized prompt with HTTP 400 and this body:

{"error":{"code":400,"message":"request (600000 tokens) exceeds the available context size (524288 tokens), try increasing it","type":"exceed_context_size_error"}}

Backend::is_context_overflow accepts an OpenAI-shaped 400 when error.code == "context_length_exceeded" or when the message contains one of OPENAI_OVERFLOW_PHRASES. llama-server does neither. Its code is the HTTP status as a number, the reason lives in type, and the message wording matches no phrase, so client.rs returned the 400 to the caller instead of moving to the next candidate and routing_fallbacks.context_window stayed at 0.

So the structured check now also accepts error.type == "exceed_context_size_error". That is the stable signal, because ERROR_TYPE_EXCEED_CONTEXT_SIZE carries two different messages in llama.cpp and the type is the same for both:

  • tools/server/server-context.cpp, "request (%d tokens) exceeds the available context size (%d tokens), try increasing it"
  • same file, "input (%d tokens) is larger than the max context size (%d tokens). skipping"

and tools/server/server-common.cpp maps the enum to exceed_context_size_error with code 400.

Both messages also go into OPENAI_OVERFLOW_PHRASES, as exceeds the available context size and larger than the max context size. That is not redundant: a proxy in front of llama-server can rebuild the error envelope and drop type while keeping the message, which is the same situation the LiteLLM entries in that list already cover.

Done as a one-concern change on main rather than waiting for #719, which the issue asks about.

Tests: openai_detects_llama_server_overflow next to the existing detection test, covering both real bodies, the type on its own with an unrelated message, the message on its own without the type, and a llama-server 400 that is not an overflow (invalid_request_error), which must still reach the caller. It fails on main on the first assert.

cargo test --workspace is 924 passed, 0 failed. cargo fmt --all --check and cargo clippy --workspace --all-targets -- -D warnings are both clean on toolchain 1.99.

Not run: nothing against a live llama-server. The bodies in the test are the ones from the issue and from the llama.cpp sources above.

Summary by CodeRabbit

  • Bug Fixes
    • Improved detection of context-size overflow errors from llama-server responses, while avoiding misclassifying unrelated errors.

llama-server rejects an oversized prompt with HTTP 400 and

    {"error":{"code":400,
               "message":"request (600000 tokens) exceeds the available context size (524288 tokens), try increasing it",
               "type":"exceed_context_size_error"}}

is_context_overflow looks for error.code == "context_length_exceeded" or one
of OPENAI_OVERFLOW_PHRASES. llama-server puts the HTTP status in code as a
number and names the reason in type, and neither of its two messages matches a
phrase, so the 400 went back to the client and the model group never fell back.

Match error.type == "exceed_context_size_error", which covers both messages
ERROR_TYPE_EXCEED_CONTEXT_SIZE can carry (tools/server/server-context.cpp),
and add both phrases for a proxy that rewrites the envelope and keeps the
message.

Closes NVIDIA-NeMo#937

Signed-off-by: Dmitry Voropaev <dy.voropaev@gmail.com>
@coderabbitai

coderabbitai Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: a142b51a-3ab5-4808-8f3d-1a8fb9fb37f2
📥 Commits

Reviewing files that changed from the base of the PR and between a3cdc51 and bfde40c.

📒 Files selected for processing (1)
  • crates/libsy-llm-client/src/backend.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

OpenAI overflow detection now recognizes two llama-server message phrases and the exceed_context_size_error error type. Tests cover these signals and confirm that an unrelated invalid_request_error is not classified as overflow.

Changes

Context overflow detection

Layer / File(s) Summary
Classify llama-server overflow responses
crates/libsy-llm-client/src/backend.rs
OpenAI overflow detection now checks two additional message phrases and error.type == "exceed_context_size_error", alongside the existing string-code check. Tests cover the error type, both phrases, and an unrelated error.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~8 minutes

Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to bfde4

The new llama.cpp overflow signals reach the existing model fallback, while unrelated 400 errors are still returned as upstream errors. No actionable merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: treating llama.cpp context-size errors as overflow.
Linked Issues check ✅ Passed [#937] The OpenAI overflow classifier now accepts error.type == "exceed_context_size_error" and both llama-server context-size message phrases. Tests cover both response bodies, each signal without …
Out of Scope Changes check ✅ Passed The changes are limited to llama-server overflow recognition and focused regression tests. Both support the fallback fix requested by [#937].
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 4 functions across 1 files.
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checked the server’s reply,
“Context too full,” it heard it cry.
Two phrases joined the type in view,
Tests checked the signals through and through.
Now onward hops the fallback route,
With lettuce tucked beneath its snout.

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[bug] Context-overflow fallback does not recognize llama.cpp's "exceeds the available context size" error

1 participant