Skip to content

[None][fix] KV cache manager V2: keep the KDA replay state across a suspend; refuse a pool rebalance with it - #19846

Draft
vsabavat wants to merge 2 commits into
NVIDIA:mainfrom
vsabavat:k3-up/p12-kda-replay-suspend
Draft

vsabavat wants to merge 2 commits into
NVIDIA:mainfrom
vsabavat:k3-up/p12-kda-replay-suspend

Conversation

@vsabavat

@vsabavat vsabavat commented Oct 4, 2026

Copy link
Copy Markdown

Description

MambaHybridCacheManagerV2 keeps the KDA replay state per SSM slot, outside the V2 pools:
prev_num_accepted_tokens and the six replay caches (kda_conv_q / kda_conv_k / kda_conv_v, kda_qkg_cache,
kda_v_cache, kda_beta_cache), allocated once from the SSM pool's slot count at construction. It remembers each
request's last slot in _request_id_to_state_index until free_resources, and _setup_state_indices relocates a
request's replay state when its slot changes. Two paths break the assumption that a request's slot holds its state.

1. A suspended request resumes with another request's replay state

A V2 suspend does not keep that state. The V2 scheduler suspends generation requests under page pressure (eviction
and self-eviction), and the pool rebalance suspends every request. A suspend only unpins the pages; with the host tier
KVCacheManagerV2 provisions by default, they move to host when another cache needs the memory, and the request's
SSM slot goes to another request. The suspended request then resumes in whichever slot is free:

  • in a new slot, _setup_state_indices relocated the old slot's contents, by then the other request's replay state;
  • back in its old slot (after the other request finished), nothing was relocated or reset.

Either way the next verify replayed another request's drafts on this request's state: wrong tokens, no error.

The fix (commit 1):

  • suspend_request (override in MambaHybridCacheManagerV2) copies the slot's replay state to host memory
    (pinned where preferred, ordered after the last step on the stream) and forgets the slot. Every suspend path goes
    through it.
  • _setup_state_indices writes a resumed request's saved state into the slot it resumes in (no relocation, no
    reset); a request that comes back as a context request drops it.
  • update_resources (host drafter): the acceptance of a batch is recorded after the next batch is scheduled. For
    a request suspended in between it now goes with the saved state instead of the old slot, which may belong to
    another request by then.
  • free_resources and shutdown drop saved state.
  • _relocate_kda_replay_slots and the copy share one list of the per-slot buffers (_kda_replay_state_buffers).

Cost: only when a KDA replay request is suspended: one device-to-host copy of its slot's replay state, held in host
memory while it is suspended, and one host-to-device copy when it resumes.

2. A KV pool rebalance can grow the SSM pool past the replay buffers

With kv_cache_config.enable_kv_pool_rebalance, the executor's rebalance calls adjust(), which resizes the pool
groups to the auto-tuner's target ratio and can give the SSM pool more slots than it had at construction. A request in
a slot past the buffers makes the prev_num_accepted_tokens[slot] = 0 reset fail, and the KDA verify kernels index
the replay entries by slot without a bound. Nothing refused the combination: _util.py refuses a rebalance only with
block reuse, and the executor's rebalance gate skips only two-model drafters, so one-model speculative decoding on
KDA layers reaches it.

The fix (commit 2): _build_cache_config raises ValueError for the KDA replay path with
enable_kv_pool_rebalance, as KvCacheConfig does for fp8_ds_mla's fixed pool views. The rebalance stays allowed
without the replay caches.

Note for #19817

#19817 (per-token KDA verify states) adds kda_state_tok, another per-slot buffer of this scheme, which it appends in
_relocate_kda_replay_slots. Whichever of the two merges second moves that append into _kda_replay_state_buffers
(two lines), so the per-token states are relocated and kept across a suspend with the rest. The rebalance refusal
keys on the replay path, so it already covers them.

Test Coverage

tests/unittest/_torch/executor/kv_cache/test_mamba_cache_manager.py:

  • test_v2_kda_replay_state_follows_a_suspended_request[new_slot|old_slot] (CPU): a request is suspended, its
    slot goes to another request, and it resumes in a new slot or in its old one; it keeps its own replay state.
    Fails before the fix in both cases (it resumes with the other request's).
  • test_v2_kda_replay_late_acceptance_follows_a_suspended_request (CPU): the host drafter's acceptance recorded
    after the request's suspend lands in its saved state, not in its old slot. Fails before the fix (the old slot,
    another request's by then, is overwritten).
  • test_v2_kda_replay_state_follows_a_suspended_request_on_the_runtime (one GPU): the same on the V2 runtime.
    The manager has Kimi K3's per-rank KDA shapes, so its SSM pool holds 5 slots. New caches fill it; the runtime
    moves the suspended request's page to the host tier and gives its slot to the fifth cache; the request resumes in
    another slot. Before the fix it resumes with that cache's replay state; after it, with its own.
  • test_v2_kda_replay_host_drafter_records_active_requests: its manager stub gains the two new attributes.
  • test_v2_kda_replay_refuses_kv_pool_rebalance[kda_replay|no_replay] (one GPU, commit 2): building the manager
    with the replay caches and enable_kv_pool_rebalance raises; without the replay caches the rebalance stays allowed.
    The first case fails before the fix (the manager builds).

Repro of the rebalance overrun on the V2 runtime (the manager at Kimi K3's shapes, no fix needed to show it): with the
auto-tuner's target giving the SSM pool group 8x its share, adjust() grows the SSM pool from 10 to 11 slots over
10-slot replay buffers.

Results on one GB200 GPU, against main at ace058d:

  • The 7 tests above on main's code: 5 fail (both state_follows_a_suspended_request cases, the late acceptance, the
    runtime test and refuses_kv_pool_rebalance[kda_replay]) and 2 pass. With this PR all 7 pass.
  • The whole file: main 34 failed / 151 passed, this PR 34 failed / 157 passed. The same 34 fail on both, and main
    waives all of them (nvbugs/6818710): get_kv_cache_manager_cls reads model_config.mapping, which their config
    stubs lack. None of them runs code this PR changes. Nothing fails only with this PR.

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title. (No API change.)

  • Any new dependencies have been scanned for license and vulnerabilities (No new dependencies.)

  • CODEOWNERS updated if ownership changes (No ownership change.)

  • Documentation updated as needed (No user-facing change.)

  • Update tava architecture diagram if there is a significant design change in PR. (No design change.)

  • The reviewers assigned automatically/manually are appropriate for the PR. (To check once the PR is open.)

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

Vasanth Sabavat added 2 commits October 3, 2026 18:17
…y state

MambaHybridCacheManagerV2 keeps the KDA replay state
(prev_num_accepted_tokens and the six replay caches) per SSM slot,
outside the V2 pools, and remembers each request's last slot until
free_resources. A V2 suspend (the scheduler's eviction and
self-eviction, the pool rebalance) only suspends the pages. Under
pressure the host tier takes them, the slot goes to another request,
and the suspended request resumes in whichever slot is free. Its next
_setup_state_indices then relocated the old slot's contents, by then
the other request's replay state, or, back in its old slot, kept them:
the next verify replayed another request's drafts.

suspend_request now copies the slot's replay state to host memory and
forgets the slot; _setup_state_indices writes it into the slot the
request resumes in. update_resources gives the host drafter's
acceptance of a request suspended after its batch was scheduled to the
saved state instead of the old slot. Relocation and the copy share one
list of the per-slot buffers.

The new tests resume a request in a new slot and in its old one and
record a late acceptance on a manager stub, and on the V2 runtime let
new caches take a suspended request's slot through the host tier before
it resumes.

Signed-off-by: Vasanth Sabavat <vsabavat@nvidia.com>
…KDA replay caches

The KDA replay caches and prev_num_accepted_tokens are allocated once,
one entry per slot the SSM pool keeps at construction
(get_page_index_upper_bound). With enable_kv_pool_rebalance, adjust()
resizes the pool groups to the auto-tuner's target ratio, which gives
the SSM pool one slot per typical request, so a pool built smaller
than that grows past the buffers. A request in a slot past them makes
the prev_num_accepted_tokens[slot] = 0 reset fail, and the KDA verify
kernels index the replay entries by slot without a bound. Nothing
refused the combination: _util.py refuses a rebalance only with block
reuse, and the executor's rebalance gate skips only two-model
drafters, so one-model speculative decoding on KDA layers reaches it.

_build_cache_config now raises ValueError for the KDA replay path with
enable_kv_pool_rebalance, as KvCacheConfig does for fp8_ds_mla's fixed
pool views. test_v2_kda_replay_refuses_kv_pool_rebalance covers the
refusal and the control without the replay caches.

Signed-off-by: Vasanth Sabavat <vsabavat@nvidia.com>
@svc-trtllm-gh-bot svc-trtllm-gh-bot added the Community want to contribute PRs initiated from Community label Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Community want to contribute PRs initiated from Community

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants