Conversation
added 2 commits
October 3, 2026 18:17
…y state MambaHybridCacheManagerV2 keeps the KDA replay state (prev_num_accepted_tokens and the six replay caches) per SSM slot, outside the V2 pools, and remembers each request's last slot until free_resources. A V2 suspend (the scheduler's eviction and self-eviction, the pool rebalance) only suspends the pages. Under pressure the host tier takes them, the slot goes to another request, and the suspended request resumes in whichever slot is free. Its next _setup_state_indices then relocated the old slot's contents, by then the other request's replay state, or, back in its old slot, kept them: the next verify replayed another request's drafts. suspend_request now copies the slot's replay state to host memory and forgets the slot; _setup_state_indices writes it into the slot the request resumes in. update_resources gives the host drafter's acceptance of a request suspended after its batch was scheduled to the saved state instead of the old slot. Relocation and the copy share one list of the per-slot buffers. The new tests resume a request in a new slot and in its old one and record a late acceptance on a manager stub, and on the V2 runtime let new caches take a suspended request's slot through the host tier before it resumes. Signed-off-by: Vasanth Sabavat <vsabavat@nvidia.com>
…KDA replay caches The KDA replay caches and prev_num_accepted_tokens are allocated once, one entry per slot the SSM pool keeps at construction (get_page_index_upper_bound). With enable_kv_pool_rebalance, adjust() resizes the pool groups to the auto-tuner's target ratio, which gives the SSM pool one slot per typical request, so a pool built smaller than that grows past the buffers. A request in a slot past them makes the prev_num_accepted_tokens[slot] = 0 reset fail, and the KDA verify kernels index the replay entries by slot without a bound. Nothing refused the combination: _util.py refuses a rebalance only with block reuse, and the executor's rebalance gate skips only two-model drafters, so one-model speculative decoding on KDA layers reaches it. _build_cache_config now raises ValueError for the KDA replay path with enable_kv_pool_rebalance, as KvCacheConfig does for fp8_ds_mla's fixed pool views. test_v2_kda_replay_refuses_kv_pool_rebalance covers the refusal and the control without the replay caches. Signed-off-by: Vasanth Sabavat <vsabavat@nvidia.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
MambaHybridCacheManagerV2keeps the KDA replay state per SSM slot, outside the V2 pools:prev_num_accepted_tokensand the six replay caches (kda_conv_q/kda_conv_k/kda_conv_v,kda_qkg_cache,kda_v_cache,kda_beta_cache), allocated once from the SSM pool's slot count at construction. It remembers eachrequest's last slot in
_request_id_to_state_indexuntilfree_resources, and_setup_state_indicesrelocates arequest's replay state when its slot changes. Two paths break the assumption that a request's slot holds its state.
1. A suspended request resumes with another request's replay state
A V2 suspend does not keep that state. The V2 scheduler suspends generation requests under page pressure (eviction
and self-eviction), and the pool rebalance suspends every request. A suspend only unpins the pages; with the host tier
KVCacheManagerV2provisions by default, they move to host when another cache needs the memory, and the request'sSSM slot goes to another request. The suspended request then resumes in whichever slot is free:
_setup_state_indicesrelocated the old slot's contents, by then the other request's replay state;Either way the next verify replayed another request's drafts on this request's state: wrong tokens, no error.
The fix (commit 1):
suspend_request(override inMambaHybridCacheManagerV2) copies the slot's replay state to host memory(pinned where preferred, ordered after the last step on the stream) and forgets the slot. Every suspend path goes
through it.
_setup_state_indiceswrites a resumed request's saved state into the slot it resumes in (no relocation, noreset); a request that comes back as a context request drops it.
update_resources(host drafter): the acceptance of a batch is recorded after the next batch is scheduled. Fora request suspended in between it now goes with the saved state instead of the old slot, which may belong to
another request by then.
free_resourcesandshutdowndrop saved state._relocate_kda_replay_slotsand the copy share one list of the per-slot buffers (_kda_replay_state_buffers).Cost: only when a KDA replay request is suspended: one device-to-host copy of its slot's replay state, held in host
memory while it is suspended, and one host-to-device copy when it resumes.
2. A KV pool rebalance can grow the SSM pool past the replay buffers
With
kv_cache_config.enable_kv_pool_rebalance, the executor's rebalance callsadjust(), which resizes the poolgroups to the auto-tuner's target ratio and can give the SSM pool more slots than it had at construction. A request in
a slot past the buffers makes the
prev_num_accepted_tokens[slot] = 0reset fail, and the KDA verify kernels indexthe replay entries by slot without a bound. Nothing refused the combination:
_util.pyrefuses a rebalance only withblock reuse, and the executor's rebalance gate skips only two-model drafters, so one-model speculative decoding on
KDA layers reaches it.
The fix (commit 2):
_build_cache_configraisesValueErrorfor the KDA replay path withenable_kv_pool_rebalance, asKvCacheConfigdoes forfp8_ds_mla's fixed pool views. The rebalance stays allowedwithout the replay caches.
Note for #19817
#19817 (per-token KDA verify states) adds
kda_state_tok, another per-slot buffer of this scheme, which it appends in_relocate_kda_replay_slots. Whichever of the two merges second moves that append into_kda_replay_state_buffers(two lines), so the per-token states are relocated and kept across a suspend with the rest. The rebalance refusal
keys on the replay path, so it already covers them.
Test Coverage
tests/unittest/_torch/executor/kv_cache/test_mamba_cache_manager.py:test_v2_kda_replay_state_follows_a_suspended_request[new_slot|old_slot](CPU): a request is suspended, itsslot goes to another request, and it resumes in a new slot or in its old one; it keeps its own replay state.
Fails before the fix in both cases (it resumes with the other request's).
test_v2_kda_replay_late_acceptance_follows_a_suspended_request(CPU): the host drafter's acceptance recordedafter the request's suspend lands in its saved state, not in its old slot. Fails before the fix (the old slot,
another request's by then, is overwritten).
test_v2_kda_replay_state_follows_a_suspended_request_on_the_runtime(one GPU): the same on the V2 runtime.The manager has Kimi K3's per-rank KDA shapes, so its SSM pool holds 5 slots. New caches fill it; the runtime
moves the suspended request's page to the host tier and gives its slot to the fifth cache; the request resumes in
another slot. Before the fix it resumes with that cache's replay state; after it, with its own.
test_v2_kda_replay_host_drafter_records_active_requests: its manager stub gains the two new attributes.test_v2_kda_replay_refuses_kv_pool_rebalance[kda_replay|no_replay](one GPU, commit 2): building the managerwith the replay caches and
enable_kv_pool_rebalanceraises; without the replay caches the rebalance stays allowed.The first case fails before the fix (the manager builds).
Repro of the rebalance overrun on the V2 runtime (the manager at Kimi K3's shapes, no fix needed to show it): with the
auto-tuner's target giving the SSM pool group 8x its share,
adjust()grows the SSM pool from 10 to 11 slots over10-slot replay buffers.
Results on one GB200 GPU, against main at ace058d:
state_follows_a_suspended_requestcases, the late acceptance, theruntime test and
refuses_kv_pool_rebalance[kda_replay]) and 2 pass. With this PR all 7 pass.waives all of them (nvbugs/6818710):
get_kv_cache_manager_clsreadsmodel_config.mapping, which their configstubs lack. None of them runs code this PR changes. Nothing fails only with this PR.
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title. (No API change.)Any new dependencies have been scanned for license and vulnerabilities (No new dependencies.)
CODEOWNERS updated if ownership changes (No ownership change.)
Documentation updated as needed (No user-facing change.)
Update tava architecture diagram if there is a significant design change in PR. (No design change.)
The reviewers assigned automatically/manually are appropriate for the PR. (To check once the PR is open.)
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.