Automated integration-failure triage for the daily CI window 2026-09-14 → 2026-09-20.
This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).
The Task column links Serge's own page for the run — the steps it went through, what each tool cost, and the full task text. 🔒 means it is not public: the host is VPN-internal and the page needs a Serge login. A task links once Serge has accepted it, so the column fills in on the first refresh after dispatch.
Dispatched failure groups
| Model |
Error |
Occurrences |
PR |
Task |
qwen3_omni_moe (regressed by PR #48714) |
mixed — other (2) |
2 |
#48957 |
🔒 task |
generation |
cuda_runtime — other (12) |
12 |
🚫 no fix |
🔒 task |
diffllama |
output_mismatch — list output differs (2) |
2 |
🚫 no fix |
🔒 task |
kosmos2 |
other — other (2) |
2 |
🚫 no fix |
🔒 task |
qwen3_vl_moe |
OOM — other (1) |
1 |
🚫 no fix |
🔒 task |
Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3).
2 group(s) skipped — stopped failing before this run (afmoe, deepseek_v3). Not dispatched and not awaiting a human.
Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
| Group |
Reason |
LLM |
Tokens (in / out) |
Failing tests |
generation |
[not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_validate_assistant · test_matches_immediate_check_at_max_length · test_matches_immediate_check_on_a_batch · +3 more |
diffllama |
DiffLlamaIntegrationTest::test_compile_static_cache fails because the model now emits degenerate repetitive text ("2.5 5 5 5 ..." and "a a a a ...") instead of the previously coherent completion. This is not a small hardware drift or a plausible variant of the expected text — it is a real regression in the DiffLlama generation path when cache_implementation="static" is used with torch.compile. |
moonshotai/Kimi-K2.7-Code |
1,857,894 / 4,361 |
test_compile_static_cache |
kosmos2 |
The test_snowman_image_captioning integration test fails because the model now generates a different caption/scores than the hard-coded expected values. The traceback shows the generated text is coherent but different ("a snowman wearing a hat and standing in front of a fire" vs the old "a snowman warming himself by a fire"), so this is a stale-expectations issue rather than a crash or library bug. |
moonshotai/Kimi-K2.7-Code |
741,940 / 9,088 |
test_snowman_image_captioning |
qwen3_vl_moe |
[not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. |
moonshotai/Kimi-K2.7-Code |
642,658 / 12,053 |
test_small_model_integration_test |
Generated 2026-09-21T00:42:16+00:00.
Automated integration-failure triage for the daily CI window
2026-09-14 → 2026-09-20.This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a
serge/fix/itf-<fingerprint>branch. This table is refreshed in place as Serge runs: a group links its#<pr>when opened, shows🚫 no fixwhen Serge found no safe change,⚠️ task failedon error, or(pending)while still running (a late PR links on the next nightly run).The Task column links Serge's own page for the run — the steps it went through, what each tool cost, and the full task text. 🔒 means it is not public: the host is VPN-internal and the page needs a Serge login. A task links once Serge has accepted it, so the column fills in on the first refresh after dispatch.
Dispatched failure groups
qwen3_omni_moe(regressed by PR #48714)generationdiffllamakosmos2qwen3_vl_moeNot dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched:
bamba(4),llama4(3).2 group(s) skipped — stopped failing before this run (
afmoe,deepseek_v3). Not dispatched and not awaiting a human.Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
generationnot_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR.moonshotai/Kimi-K2.7-CodediffllamaDiffLlamaIntegrationTest::test_compile_static_cachefails because the model now emits degenerate repetitive text ("2.5 5 5 5 ..."and"a a a a ...") instead of the previously coherent completion. This is not a small hardware drift or a plausible variant of the expected text — it is a real regression in the DiffLlama generation path whencache_implementation="static"is used withtorch.compile.moonshotai/Kimi-K2.7-Codekosmos2test_snowman_image_captioningintegration test fails because the model now generates a different caption/scores than the hard-coded expected values. The traceback shows the generated text is coherent but different ("a snowman wearing a hat and standing in front of a fire"vs the old"a snowman warming himself by a fire"), so this is a stale-expectations issue rather than a crash or library bug.moonshotai/Kimi-K2.7-Codeqwen3_vl_moenot_fixed): the patch did not turn the targeted tests green.moonshotai/Kimi-K2.7-CodeGenerated 2026-09-21T00:42:16+00:00.