Repository navigation
Conversation
llama-server rejects an oversized prompt with HTTP 400 and
{"error":{"code":400,
"message":"request (600000 tokens) exceeds the available context size (524288 tokens), try increasing it",
"type":"exceed_context_size_error"}}
is_context_overflow looks for error.code == "context_length_exceeded" or one
of OPENAI_OVERFLOW_PHRASES. llama-server puts the HTTP status in code as a
number and names the reason in type, and neither of its two messages matches a
phrase, so the 400 went back to the client and the model group never fell back.
Match error.type == "exceed_context_size_error", which covers both messages
ERROR_TYPE_EXCEED_CONTEXT_SIZE can carry (tools/server/server-context.cpp),
and add both phrases for a proxy that rewrites the envelope and keeps the
message.
Closes NVIDIA-NeMo#937
Signed-off-by: Dmitry Voropaev <dy.voropaev@gmail.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughOpenAI overflow detection now recognizes two llama-server message phrases and the ChangesContext overflow detection
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~8 minutes Severity of issue fixed: Medium Merge Risk: ⚪ Minimal · up to The new llama.cpp overflow signals reach the existing model fallback, while unrelated 400 errors are still returned as upstream errors. No actionable merge-blocking risk remains. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
A rabbit checked the server’s reply, Comment |
Closes #937.
llama-server answers an oversized prompt with HTTP 400 and this body:
{"error":{"code":400,"message":"request (600000 tokens) exceeds the available context size (524288 tokens), try increasing it","type":"exceed_context_size_error"}}Backend::is_context_overflowaccepts an OpenAI-shaped 400 whenerror.code == "context_length_exceeded"or when the message contains one ofOPENAI_OVERFLOW_PHRASES. llama-server does neither. Itscodeis the HTTP status as a number, the reason lives intype, and the message wording matches no phrase, soclient.rsreturned the 400 to the caller instead of moving to the next candidate androuting_fallbacks.context_windowstayed at 0.So the structured check now also accepts
error.type == "exceed_context_size_error". That is the stable signal, becauseERROR_TYPE_EXCEED_CONTEXT_SIZEcarries two different messages in llama.cpp and the type is the same for both:tools/server/server-context.cpp,"request (%d tokens) exceeds the available context size (%d tokens), try increasing it""input (%d tokens) is larger than the max context size (%d tokens). skipping"and
tools/server/server-common.cppmaps the enum toexceed_context_size_errorwith code 400.Both messages also go into
OPENAI_OVERFLOW_PHRASES, asexceeds the available context sizeandlarger than the max context size. That is not redundant: a proxy in front of llama-server can rebuild the error envelope and droptypewhile keeping the message, which is the same situation the LiteLLM entries in that list already cover.Done as a one-concern change on
mainrather than waiting for #719, which the issue asks about.Tests:
openai_detects_llama_server_overflownext to the existing detection test, covering both real bodies, the type on its own with an unrelated message, the message on its own without the type, and a llama-server 400 that is not an overflow (invalid_request_error), which must still reach the caller. It fails onmainon the first assert.cargo test --workspaceis 924 passed, 0 failed.cargo fmt --all --checkandcargo clippy --workspace --all-targets -- -D warningsare both clean on toolchain 1.99.Not run: nothing against a live llama-server. The bodies in the test are the ones from the issue and from the llama.cpp sources above.
Summary by CodeRabbit