Conversation
LauraGPT
approved these changes
Sep 26, 2026
LauraGPT
left a comment
Collaborator
There was a problem hiding this comment.
Verified exact head e2dbd93 against current main 41778c4. The loader only reads the selected source state before load_state_dict; focused repository tests pass 13/13, independent direct/wrapped/DDP/mismatch contract probes pass 6/6 on both base and head, and a separate-process 128 MiB checkpoint probe reduces peak RSS by about 143 MiB while preserving loaded values. The merge tree is conflict-free; compileall and diff-check pass.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related to #3728
Summary
load_pretrained_model()deep-copies the whole checkpoint state dict on every load. That copy is redundant:ori_stateis created bytorch.loadinside the function, so it has a single owner and is not shared with the caller.ori_stateis not referenced again after thedeepcopyline.src_stateis only read from that point on. The loop inspectssrc_state[k_src].shapeand rebinds references indst_state, andload_state_dictis what copies data into the model. No tensor is mutated in place.So the deep copy changes nothing while holding a second full copy of the checkpoint in memory for the whole load. Peak host memory during load drops by roughly the size of the checkpoint: for
iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online(220M params, 956 tensors, 840MB checkpoint) peak RSS goes from 3461 MB to 2624 MB, a 837MB saving that matches the checkpoint size. That matters on memory-constrained hosts and under container memory limits, where the current peak can fail a load that would otherwise fit.The now-unused
import copyis removed as well.Type of change
Validation
python -m compileall funasr examples testsSame machine, same checkpoint, alternating runs, via
AutoModel(..., device="cpu", disable_update=True):The ~840MB reduction is stable across runs and tracks the checkpoint size, which is the copy being dropped.
All keys matched successfullyon both sides is the correctness signal: the same parameters are loaded either way.Wall-clock is not quantified here: the host runs other builds and tests concurrently, so the load time varies with unrelated load. Directionally the copy is also faster to skip, but treat the memory figure as the reproducible result.
Runtime check on this branch after the change, streaming ASR at 600ms chunks on an RTX 4060:
User impact
Anyone loading a FunASR model — particularly on hosts with a tight memory budget, in containers with memory limits, or on the edge deployments the toolkit targets. The saving scales with checkpoint size, so it is most visible on the large ASR and LLM-ASR checkpoints.
Notes for reviewers
ori_statehas exactly one owner and is dead after the copy, andsrc_stateis read-only downstream.ori_stateto survive the call, that would be a different contract. It is not the case today, and the variable is local to this function.resource.getrusage-based rather than a unit test, so it is reported as a benchmark rather than a CI assertion.