Skip to content

mmap model weights instead of copying them into anonymous memory - #195

Open
montvid wants to merge 1 commit into
handy-computer:mainfrom
montvid:mmap-weights
Open

montvid wants to merge 1 commit into
handy-computer:mainfrom
montvid:mmap-weights

Conversation

@montvid

@montvid montvid commented Oct 5, 2026

Copy link
Copy Markdown

Problem

Weights are read tensor-by-tensor with fin.read() into a staging vector and copied in with ggml_backend_tensor_set (transcribe-load-common.cpp), so the entire model lives in anonymous memory. grep -i mmap src/ currently returns nothing.

Anonymous pages can't be reclaimed, so Android's low-memory killer charges every weight byte to the process. On a 5.5 GB phone I measured lmkd: Reclaim 'dev.notune.transcribe' ... to free 1531868kB rss; reason: min2x watermark is breached even after kill — a 1.7B Qwen3-ASR would load and transcribe a one-off file fine, then get killed as soon as a second component (the voice-input panel) was resident at the same time.

This is orthogonal to #150/#161/#162, which target compute/scratch growth. This one is purely about where the weights live. #162's note that "Qwen3-ASR 0.6B — ~7.6 GiB RSS ... for an ~811 MiB model" is a good illustration: this patch removes that 811 MiB component from the anonymous total; the rest is still scratch.

Approach

load_common::map_tensor_data_cpu() mmaps the GGUF PROT_READ/MAP_PRIVATE, madvise(MADV_RANDOM)s the tensor region, wraps it in ggml_backend_cpu_buffer_from_ptr (whose free_buffer is NULL, so it doesn't own the pages) and points each tensor at its file offset with ggml_backend_tensor_alloc. A MappedWeights RAII member on the model owns the mapping and is released after the buffer and ctx_meta, both of which point into it. Same shape as llama.cpp's mmap path.

It returns bool, not transcribe_status, on purpose: mapping is an optimization, and there's no suitable "not supported, carry on" code in the enum. Every failure — Windows, non-CPU primary, open/mmap failure, or a file whose tensors are misaligned or run past EOF — falls back to the existing alloc_ctx_tensors + stream_tensor_data path, which re-validates and reports genuine corruption itself. All tensors are checked before any is allocated, so a bad file falls back cleanly instead of half-mapping ctx_meta.

Safety

Mapping read-only turns a post-load write to a weight tensor into a crash, so each arch was audited for ggml_backend_tensor_set targets before wiring:

  • qwen3_asr — all targets are graph-input tensors in their own contexts; packed gate/up are new tensors.
  • parakeet — BN fusion only reads the raw BN tensors and writes into its own bn_fused buffer; the conv_pw F32 promotion emits new tensors; the decoder's targets are its own joint/LSTM buffers.
  • whisper — the GGUF path only reads frontend.mel_filterbank / frontend.window into host buffers. The legacy .bin loader does write those two in place, so only the GGUF path is mapped and bin_load.cpp keeps its own allocation.

voxtral, voxtral_realtime, granite, funasr_nano and medasr are untouched and still stream — each needs the same audit before adopting it. The doc comment on map_tensor_data_cpu says so.

Measurements

Peak RssAnon, x86 Linux, Release, -DGGML_NATIVE=ON, 4 threads, one 8.6 s clip:

model file before after
qwen3-asr 1.7B Q4_0 1114 MiB 1936 MiB 817 MiB
parakeet 0.6B Q8_0 705 MiB 1128 MiB 423 MiB
whisper distil Q8_0 792 MiB 897 MiB 102 MiB

The reduction tracks file size. Peak RSS is roughly unchanged — file pages still count toward RSS — but the anonymous share, which is what the OOM killer charges, drops by the size of the model.

On the device that motivated this (Galaxy S21 FE, Android 16), the 1.7B went from being reclaimed to running the voice-input panel with 0 lmkd kills.

Verification

Transcripts are byte-identical before and after on all three architectures. Dropping the staging copy makes load slightly faster (8.27 s → 7.38 s for the 1.7B on-device); the only cost is page-fault latency on the first run against a cold page cache — I measured 1.74 s then 1.22/1.23 s for parakeet, against 1.20 s unmapped.

Happy to split this per-arch, gate it behind an option, or drop the whisper/parakeet wiring if you'd rather land the helper alone first.

Weights are currently read tensor-by-tensor with fin.read() into a staging
vector and copied in with ggml_backend_tensor_set, so the whole model lives in
anonymous heap. Anonymous pages cannot be reclaimed, so on Android every weight
byte is charged to the process by the low-memory killer; file-backed pages can
simply be dropped and re-read. On a 5.5 GB phone this is the difference between
a 1.7B model running and being killed mid-session.

Add load_common::map_tensor_data_cpu(), which mmaps the GGUF read-only, wraps
the tensor-data region in ggml_backend_cpu_buffer_from_ptr (whose free_buffer
is NULL, so it does not own the pages) and points each tensor at its file
offset with ggml_backend_tensor_alloc. A MappedWeights RAII member on the model
owns the mapping and is released after the buffer and ctx_meta, both of which
point into it.

It returns bool rather than transcribe_status: mapping is an optimization, so
every failure - unsupported platform, non-CPU backend, open/mmap failure, or a
misaligned or truncated file - falls back to the existing alloc_ctx_tensors +
stream_tensor_data path, which re-validates and reports real corruption itself.
Every tensor is checked before any is allocated so a bad file falls back
cleanly rather than half-mapping ctx_meta.

Wired up for qwen3_asr, parakeet and whisper. Each was audited first for writes
to ctx_meta tensors after load, since mapping read-only turns such a write into
a crash. Notably whisper's legacy .bin loader does write frontend.mel_filterbank
and frontend.window in place, so only the GGUF path is mapped and bin_load.cpp
keeps its own allocation. voxtral, granite, funasr_nano and medasr are untouched
and still stream.

Peak RssAnon, x86 Linux, Release, 4 threads, one 8.6 s clip:

  model                     file     before    after
  qwen3-asr 1.7B Q4_0    1114 MiB   1936 MiB   817 MiB
  parakeet 0.6B Q8_0      705 MiB   1128 MiB   423 MiB
  whisper distil Q8_0     792 MiB    897 MiB   102 MiB

The reduction tracks file size, as expected. Peak RSS is roughly unchanged
(file pages still count) but the anonymous share, which is what the OOM killer
charges, drops by the size of the model. Transcripts are identical on all
three, and dropping the staging copy makes load slightly faster; the only cost
is page-fault latency on the first run against a cold page cache.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@montvid
montvid requested a review from cjpais as a code owner October 5, 2026 15:22
@montvid

montvid commented Oct 5, 2026

Copy link
Copy Markdown
Author

Benchmarks via transcribe-bench

Re-measured with the repo's own harness rather than my ad-hoc timing. Same machine, same commit for both builds (3727340), -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON, --iters 3 --warmup 1 --threads 4, one 8.6 s 16 kHz clip. Peak RssAnon / VmHWM sampled from /proc/<pid>/status at 100 ms while the benchmark ran.

model file base mmap
qwen3-asr 1.7B Q4_0 1114 MiB load_ms 1154.2 350.6
wall_ms mean 4209.7 4331.7
rtf_wall 2.05 2.00
peak RssAnon 1936 MiB 828 MiB
peak RSS 1943 MiB 1942 MiB
parakeet 0.6B Q8_0 705 MiB load_ms 859.4 390.3
wall_ms mean 723.9 724.5
rtf_wall 11.94 11.93
peak RssAnon 1133 MiB 428 MiB
peak RSS 1139 MiB 1139 MiB
whisper distil Q8_0 792 MiB load_ms 633.4 33.9
wall_ms mean 11081.5 11193.1
rtf_wall 0.78 0.77
peak RssAnon 916 MiB 114 MiB
peak RSS 923 MiB 908 MiB

Output is unchanged. hyp_text is byte-identical between the two builds for all three models (sha256 of the hypothesis matches), and for parakeet — the one family where transcribe-bench populates token_ids_csv — all 39 token IDs are identical too. qwen3_asr and whisper report n_tokens: 0 in the bench JSON so there were no IDs to compare; their hypothesis strings match exactly.

Load gets faster, which I had not expected going in: 18.7× for whisper (633 → 34 ms), 3.3× for qwen (1154 → 351 ms), 2.2× for parakeet (859 → 390 ms). Mapping replaces a full read-plus-copy of the file with page-table setup, and pages fault in lazily during the first run.

rtf_wall is unchanged in every case (2.05→2.00, 11.94→11.93, 0.78→0.77 — all within run-to-run noise). Peak RSS is also unchanged, as expected: the mapped pages still count toward RSS. The point is the split, not the total — RssAnon drops by very close to the file size in each case, and that is what an OOM killer charges against the process.

On-device

Galaxy S21 FE (5.5 GB RAM, Android 16), measured through a host app, RssAnon in kB:

model 0.1.3 0.3.1 0.3.1 + this patch
qwen3-asr 1.7B Q4_0 2,064,316 1,786,472 700,504
parakeet 0.6B Q8_0 1,204,496 1,160,548 439,436
whisper distil Q8_0 — — 97,064

The middle column is worth noting on its own: the scratch-memory work already merged in #150/#161 is worth ~278 MB on the 1.7B between 0.1.3 and 0.3.1. This patch is additive to that, and on this device it is the difference between the 1.7B being reclaimed mid-session and running with zero lmkd kills.

Warm-run latency on-device is unchanged (parakeet 1.24 s vs 1.20–1.29 s before); only the first run against a cold page cache costs ~0.7 s extra.

The patch applies unmodified to the released transcribe-cpp-sys 0.3.1 crate as well as to main.

@cjpais

cjpais commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Thanks we definitely should do this

@montvid

montvid commented Oct 8, 2026

Copy link
Copy Markdown
Author

Thanks, my phone has 6 gb ram and 6 gb swap, no other way 1.5 gb asr models fit in it only with mmap.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants