Skip to content

[serge] integration failure triage - 2026-09-20 #48968

Description

@github-actions

Automated integration-failure triage for the daily CI window 2026-09-14 → 2026-09-20.

This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.

Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).

The Task column links Serge's own page for the run — the steps it went through, what each tool cost, and the full task text. 🔒 means it is not public: the host is VPN-internal and the page needs a Serge login. A task links once Serge has accepted it, so the column fills in on the first refresh after dispatch.

Dispatched failure groups

Model Error Occurrences PR Task
qwen3_omni_moe (regressed by PR #48714) mixed — other (2) 2 #48957 🔒 task
generation cuda_runtime — other (12) 12 🚫 no fix 🔒 task
diffllama output_mismatch — list output differs (2) 2 🚫 no fix 🔒 task
kosmos2 other — other (2) 2 🚫 no fix 🔒 task
qwen3_vl_moe OOM — other (1) 1 🚫 no fix 🔒 task

Not dispatched — environment / dependency

These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).

2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3).

2 group(s) skipped — stopped failing before this run (afmoe, deepseek_v3). Not dispatched and not awaiting a human.

Outcome recap

Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.

Group Reason LLM Tokens (in / out) Failing tests
generation [not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. moonshotai/Kimi-K2.7-Code — / — test_validate_assistant · test_matches_immediate_check_at_max_length · test_matches_immediate_check_on_a_batch · +3 more
diffllama DiffLlamaIntegrationTest::test_compile_static_cache fails because the model now emits degenerate repetitive text ("2.5 5 5 5 ..." and "a a a a ...") instead of the previously coherent completion. This is not a small hardware drift or a plausible variant of the expected text — it is a real regression in the DiffLlama generation path when cache_implementation="static" is used with torch.compile. moonshotai/Kimi-K2.7-Code 1,857,894 / 4,361 test_compile_static_cache
kosmos2 The test_snowman_image_captioning integration test fails because the model now generates a different caption/scores than the hard-coded expected values. The traceback shows the generated text is coherent but different ("a snowman wearing a hat and standing in front of a fire" vs the old "a snowman warming himself by a fire"), so this is a stale-expectations issue rather than a crash or library bug. moonshotai/Kimi-K2.7-Code 741,940 / 9,088 test_snowman_image_captioning
qwen3_vl_moe [not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. moonshotai/Kimi-K2.7-Code 642,658 / 12,053 test_small_model_integration_test

Generated 2026-09-21T00:42:16+00:00.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions