Skip to content

[None][fix] Scope KVCM warmup capacity constraints to DeepSeek V4 - #19213

Open
yizhang-nv wants to merge 15 commits into
NVIDIA:mainfrom
yizhang-nv:codex/fix-kvcm-v2-init-warmup-budget
Open

yizhang-nv wants to merge 15 commits into
NVIDIA:mainfrom
yizhang-nv:codex/fix-kvcm-v2-init-warmup-budget

Conversation

@yizhang-nv

@yizhang-nv yizhang-nv commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

Description

DeepSeek V4's maximum-sequence-length warmup constraint was moved into generic KVCM V2 in #16545. For ordinary attention models, that hard floor can enlarge the temporary KV pool beyond its estimated GPU budget, causing OOM during cache creation or encoder profiling.

Restore the longest-decode-plus-short-requests constraint to DeepseekV4CacheManager. Keep the general context/chunked-prefill constraint in generic V2: it covers the configured per-iteration token budget. Preserve V4's draft/extra reservations, explicit pool-ratio opt-out, and generic average-length pool preferences.

Also correct generation dummy allocation at the source. token_nums already includes history plus the current input; the generation resize counted that input again. Remove only this duplicate +1, retaining draft/extra reservations and the normal scheduler's generation growth. This lets CUDA-graph warmup use the capacity returned by the manager without skipping feasible batch shapes. model_engine.py is unchanged from main.

Original gRPC, Seed-OSS, Mistral, and multimodal-example fixtures explicitly select V2. A benchmark comment is corrected without changing its budget. No native allocator or public configuration changes.

Test Coverage

The committed generic test changes retain the context constraint assertions and add one regression with draft lengths 0 and 4: query the available capacity, then allocate a generation dummy exactly at a page boundary. Both cases fail on the old dummy allocator and pass with the correction. The enclosing executor directory is already included in l0_h100.yml.

One CPU-only V4 configuration test covers the default ratio and an explicit pool ratio. It asserts the exact longest-decode-plus-minimal-decodes constraint (including draft/extra reservations), preserves the inherited context constraint, and checks the explicit-ratio opt-out. The existing l0_cpu.yml attention-directory entry collects its cpu_only marker. Diagnostic scripts and larger temporary matrices remain outside the PR.

Real B200 validation on September 22 PDT / September 23 UTC:

  • V4 config regression: 2 passed with all GPUs hidden and CUDA uninitialized, loading the complete current V4 source and tracked test file against the compatible CI60862 runtime. A direct whole-branch run was blocked at import by the older native runtime lacking StorageStatistics.
  • Native page-boundary regression: baseline 2 expected failures; fixed 2 passed.
  • Qwen3-0.6B runtime trace: normal decode with history 31 grows capacity 31 -> 32. A dummy with the same history and one input allocated 33 before, 32 after; both send KV length 32 to attention. Capture/replay succeeds and greedy outputs are identical.
  • Seed-OSS-36B: temporary and final caches each complete 34/34 graph warmups and captures, with no skipped shapes. Full 1,319-sample GSM8K passed at 92.077%, threshold 87.597%.
  • Qwen3-8B + Qwen3-0.6B DraftTarget, explicit V2 target and separate draft managers, D=4, graphs enabled: passed exact greedy parity for both eight-token outputs. All 8 additional temporary native-budget/capture checks passed (9/9 total, zero skips).
  • Ruff and repository commit hooks passed.

The Seed pytest completed successfully and workers shut down cleanly; its outer temporary shell runner subsequently exited 2 because that script was edited while the long test ran. This harness error and its correction are preserved in the evidence report; it is separate from the passing pytest/JUnit result.

These are isolated Python-policy comparisons on the matching CI60862 native/Python runtime (40466ac6c0), using original model tests frozen at eb7be9db6a; they are not a fresh native build of the rebased branch (63e64e5bdb). Earlier same-runtime validation reproduced and fixed the original A10 gRPC, B200 Seed, and H100 Mistral OOMs. A10/H100 full-model tests were not repeated for this final dummy-token correction; dynamic-tree and Helix consumers were checked in source, not rerun on hardware.

Evidence, module hashes, full stdout/stderr, and JUnit paths: tmp/v4-relocation/dummy-token-fix/RESULTS.md and results-summary.json in the PR workspace; full logs under /home/scratch.yizhan_sw_1/logs/2026-09-22/.

PR Checklist

  • Description and regression coverage reflect the final constraint ownership and dummy-token accounting.
  • Original model workloads, accuracy thresholds, and parity assertions are preserved.
  • No new dependency, public API/configuration field, ownership, or architecture change.

GitHub Bot Help

To see available CI bot commands, comment /bot help.

@yizhang-nv

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast --extra-stage "DGX_B200-PyTorch-Post-Merge-1,DGX_B200-PyTorch-Post-Merge-2,DGX_H100-PyTorch-Post-Merge-1,DGX_H100-PyTorch-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #73564 [ run ] triggered by Bot. Commit: bb4d145 Link to invocation

@yizhang-nv

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast --stage-list "A10-PyTorch-1,A10-PyTorch-2,A10-PyTorch-3,DGX_H100-PyTorch-1,DGX_H100-PyTorch-2,DGX_H100-PyTorch-3,DGX_H100-PyTorch-4,DGX_H100-PyTorch-5,DGX_H100-PyTorch-6,DGX_B200-PyTorch-Post-Merge-1,DGX_B200-PyTorch-Post-Merge-2,DGX_H100-PyTorch-Post-Merge-1,DGX_H100-PyTorch-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #73573 [ run ] triggered by Bot. Commit: cc9e661 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #73564 [ run ] completed with state ABORTED. Commit: bb4d145

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #73573 [ run ] completed with state SUCCESS. Commit: cc9e661
/LLM/main/L0_MergeRequest_PR pipeline #60451 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@yizhang-nv yizhang-nv changed the title [None][fix] Respect KVCM V2 initialization and warmup budgets [None][fix] Bound KVCM V2 initialization and query warmup capacity Sep 16, 2026
@yizhang-nv
yizhang-nv force-pushed the codex/fix-kvcm-v2-init-warmup-budget branch from 61d98aa to 72adcea Compare September 16, 2026 07:49
@yizhang-nv yizhang-nv changed the title [None][fix] Bound KVCM V2 initialization and query warmup capacity [None][fix] Respect KVCM V2 initialization and warmup budgets Sep 16, 2026
@yizhang-nv

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #73797 [ run ] triggered by Bot. Commit: ad12477 Link to invocation

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Semantic conflict review

The verdict of record is the Semantic conflict with target branch / PR #19213 commit status on the requested head commit. This summary updates on reply events and may lag between a new request and its reply.

Latest recorded state: No semantic conflict found (best effort) for head abb6079e9018913eacdbf56bd0290cb69c52f5ab, target ca37c9f725922f7e72017ff5f5b9794731915a34, merge base ee510fc85d39477e6ef99e79583b64e968317c1c (request 93b46dad-7a47-4a9e-8be6-71b479112ea4). CodeRabbit analysis.

Best-effort AI judgment for the recorded revisions. PASS, FAIL and INCONCLUSIVE may be incomplete or incorrect. PR authors and reviewers should independently verify the evidence and relevant behavior. This semantic review and its status/workflow are advisory, not required merge checks under current repository rules; other merge requirements still apply. Advisory status does not make a confirmed defect safe to ignore.

Requested (UTC) Head Target Verdict Comment
2026-10-02T16:40:33Z abb6079e9018 ca37c9f72592 PASS reply
2026-10-02T04:42:34Z abb6079e9018 8a3c90311ea9 PASS reply
2026-10-01T20:49:53Z abb6079e9018 80509acfc073 PASS reply
2026-10-01T14:42:21Z abb6079e9018 ee510fc85d39 PASS reply
2026-10-01T06:57:04Z 8341a33b9261 0d3bbd257d35 FAIL reply
2026-10-01T01:24:24Z 8341a33b9261 534e1f8ad9d5 FAIL reply
2026-09-30T16:41:50Z 8341a33b9261 324a51deff2a FAIL reply
2026-09-30T10:40:05Z 8341a33b9261 abddfc991456 INCONCLUSIVE reply
2026-09-30T02:50:11Z 8341a33b9261 111687f92f4d INCONCLUSIVE reply
2026-09-29T18:46:05Z 8341a33b9261 bcb288a21105 FAIL reply
2026-09-29T12:52:17Z 8341a33b9261 bcdc2d5aaf0c FAIL reply
2026-09-29T04:43:06Z 8341a33b9261 ae4aa5d6c932 FAIL reply
2026-09-28T22:35:11Z 8341a33b9261 27ba69c78189 FAIL reply
2026-09-28T14:42:44Z 8341a33b9261 0c58480ca680 FAIL reply
2026-09-28T08:53:08Z 8341a33b9261 c2220eef33f4 FAIL reply

Processed request and reply comments are minimized to reduce timeline noise; they remain expandable for audit.

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
…ek V4

Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Restore CUDA graph batch clipping when estimating V4's long-decode cache
constraint. Keep the generation dummy input counted once and align the
allocation tests with that contract, including the actual graph builder.

Validate dynamic draft schedules against observed batch sizes while requiring
coverage of each configured draft length, without depending on exact request
admission thresholds.

Validation: 154 cases passed on a B200 using the CI 62214 wheel with the V4
runtime fix overlaid, including 15 attention graph capture/replay cases.

Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
Signed-off-by: Yi Zhang <187001205+yizhang-nv@users.noreply.github.com>
@yizhang-nv
yizhang-nv force-pushed the codex/fix-kvcm-v2-init-warmup-budget branch from abb6079 to 45fc4ac Compare October 4, 2026 04:34
@yizhang-nv

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76185 [ run ] triggered by Bot. Commit: 45fc4ac Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76185 [ run ] completed with state SUCCESS. Commit: 45fc4ac
/LLM/main/L0_MergeRequest_PR pipeline #62810 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@yizhang-nv

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76191 [ run ] triggered by Bot. Commit: 45fc4ac Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76191 [ run ] completed with state SUCCESS. Commit: 45fc4ac
/LLM/main/L0_MergeRequest_PR pipeline #62816 completed with status: 'SUCCESS'

CI Report

Link to invocation

@yizhang-nv
yizhang-nv enabled auto-merge (squash) October 5, 2026 03:44

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.