Repository navigation
Conversation
Weights are currently read tensor-by-tensor with fin.read() into a staging vector and copied in with ggml_backend_tensor_set, so the whole model lives in anonymous heap. Anonymous pages cannot be reclaimed, so on Android every weight byte is charged to the process by the low-memory killer; file-backed pages can simply be dropped and re-read. On a 5.5 GB phone this is the difference between a 1.7B model running and being killed mid-session. Add load_common::map_tensor_data_cpu(), which mmaps the GGUF read-only, wraps the tensor-data region in ggml_backend_cpu_buffer_from_ptr (whose free_buffer is NULL, so it does not own the pages) and points each tensor at its file offset with ggml_backend_tensor_alloc. A MappedWeights RAII member on the model owns the mapping and is released after the buffer and ctx_meta, both of which point into it. It returns bool rather than transcribe_status: mapping is an optimization, so every failure - unsupported platform, non-CPU backend, open/mmap failure, or a misaligned or truncated file - falls back to the existing alloc_ctx_tensors + stream_tensor_data path, which re-validates and reports real corruption itself. Every tensor is checked before any is allocated so a bad file falls back cleanly rather than half-mapping ctx_meta. Wired up for qwen3_asr, parakeet and whisper. Each was audited first for writes to ctx_meta tensors after load, since mapping read-only turns such a write into a crash. Notably whisper's legacy .bin loader does write frontend.mel_filterbank and frontend.window in place, so only the GGUF path is mapped and bin_load.cpp keeps its own allocation. voxtral, granite, funasr_nano and medasr are untouched and still stream. Peak RssAnon, x86 Linux, Release, 4 threads, one 8.6 s clip: model file before after qwen3-asr 1.7B Q4_0 1114 MiB 1936 MiB 817 MiB parakeet 0.6B Q8_0 705 MiB 1128 MiB 423 MiB whisper distil Q8_0 792 MiB 897 MiB 102 MiB The reduction tracks file size, as expected. Peak RSS is roughly unchanged (file pages still count) but the anonymous share, which is what the OOM killer charges, drops by the size of the model. Transcripts are identical on all three, and dropping the staging copy makes load slightly faster; the only cost is page-fault latency on the first run against a cold page cache. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Benchmarks via
|
| model | file | base | mmap | |
|---|---|---|---|---|
| qwen3-asr 1.7B Q4_0 | 1114 MiB | load_ms |
1154.2 | 350.6 |
wall_ms mean |
4209.7 | 4331.7 | ||
rtf_wall |
2.05 | 2.00 | ||
| peak RssAnon | 1936 MiB | 828 MiB | ||
| peak RSS | 1943 MiB | 1942 MiB | ||
| parakeet 0.6B Q8_0 | 705 MiB | load_ms |
859.4 | 390.3 |
wall_ms mean |
723.9 | 724.5 | ||
rtf_wall |
11.94 | 11.93 | ||
| peak RssAnon | 1133 MiB | 428 MiB | ||
| peak RSS | 1139 MiB | 1139 MiB | ||
| whisper distil Q8_0 | 792 MiB | load_ms |
633.4 | 33.9 |
wall_ms mean |
11081.5 | 11193.1 | ||
rtf_wall |
0.78 | 0.77 | ||
| peak RssAnon | 916 MiB | 114 MiB | ||
| peak RSS | 923 MiB | 908 MiB |
Output is unchanged. hyp_text is byte-identical between the two builds for all three models (sha256 of the hypothesis matches), and for parakeet — the one family where transcribe-bench populates token_ids_csv — all 39 token IDs are identical too. qwen3_asr and whisper report n_tokens: 0 in the bench JSON so there were no IDs to compare; their hypothesis strings match exactly.
Load gets faster, which I had not expected going in: 18.7× for whisper (633 → 34 ms), 3.3× for qwen (1154 → 351 ms), 2.2× for parakeet (859 → 390 ms). Mapping replaces a full read-plus-copy of the file with page-table setup, and pages fault in lazily during the first run.
rtf_wall is unchanged in every case (2.05→2.00, 11.94→11.93, 0.78→0.77 — all within run-to-run noise). Peak RSS is also unchanged, as expected: the mapped pages still count toward RSS. The point is the split, not the total — RssAnon drops by very close to the file size in each case, and that is what an OOM killer charges against the process.
On-device
Galaxy S21 FE (5.5 GB RAM, Android 16), measured through a host app, RssAnon in kB:
| model | 0.1.3 | 0.3.1 | 0.3.1 + this patch |
|---|---|---|---|
| qwen3-asr 1.7B Q4_0 | 2,064,316 | 1,786,472 | 700,504 |
| parakeet 0.6B Q8_0 | 1,204,496 | 1,160,548 | 439,436 |
| whisper distil Q8_0 | — | — | 97,064 |
The middle column is worth noting on its own: the scratch-memory work already merged in #150/#161 is worth ~278 MB on the 1.7B between 0.1.3 and 0.3.1. This patch is additive to that, and on this device it is the difference between the 1.7B being reclaimed mid-session and running with zero lmkd kills.
Warm-run latency on-device is unchanged (parakeet 1.24 s vs 1.20–1.29 s before); only the first run against a cold page cache costs ~0.7 s extra.
The patch applies unmodified to the released transcribe-cpp-sys 0.3.1 crate as well as to main.
|
Thanks we definitely should do this |
|
Thanks, my phone has 6 gb ram and 6 gb swap, no other way 1.5 gb asr models fit in it only with mmap. |
Problem
Weights are read tensor-by-tensor with
fin.read()into a staging vector and copied in withggml_backend_tensor_set(transcribe-load-common.cpp), so the entire model lives in anonymous memory.grep -i mmap src/currently returns nothing.Anonymous pages can't be reclaimed, so Android's low-memory killer charges every weight byte to the process. On a 5.5 GB phone I measured
lmkd: Reclaim 'dev.notune.transcribe' ... to free 1531868kB rss; reason: min2x watermark is breached even after kill— a 1.7B Qwen3-ASR would load and transcribe a one-off file fine, then get killed as soon as a second component (the voice-input panel) was resident at the same time.This is orthogonal to #150/#161/#162, which target compute/scratch growth. This one is purely about where the weights live. #162's note that "Qwen3-ASR 0.6B — ~7.6 GiB RSS ... for an ~811 MiB model" is a good illustration: this patch removes that 811 MiB component from the anonymous total; the rest is still scratch.
Approach
load_common::map_tensor_data_cpu()mmaps the GGUFPROT_READ/MAP_PRIVATE,madvise(MADV_RANDOM)s the tensor region, wraps it inggml_backend_cpu_buffer_from_ptr(whosefree_bufferisNULL, so it doesn't own the pages) and points each tensor at its file offset withggml_backend_tensor_alloc. AMappedWeightsRAII member on the model owns the mapping and is released after the buffer andctx_meta, both of which point into it. Same shape as llama.cpp's mmap path.It returns
bool, nottranscribe_status, on purpose: mapping is an optimization, and there's no suitable "not supported, carry on" code in the enum. Every failure — Windows, non-CPU primary,open/mmapfailure, or a file whose tensors are misaligned or run past EOF — falls back to the existingalloc_ctx_tensors+stream_tensor_datapath, which re-validates and reports genuine corruption itself. All tensors are checked before any is allocated, so a bad file falls back cleanly instead of half-mappingctx_meta.Safety
Mapping read-only turns a post-load write to a weight tensor into a crash, so each arch was audited for
ggml_backend_tensor_settargets before wiring:bn_fusedbuffer; the conv_pw F32 promotion emits new tensors; the decoder's targets are its own joint/LSTM buffers.frontend.mel_filterbank/frontend.windowinto host buffers. The legacy.binloader does write those two in place, so only the GGUF path is mapped andbin_load.cppkeeps its own allocation.voxtral,voxtral_realtime,granite,funasr_nanoandmedasrare untouched and still stream — each needs the same audit before adopting it. The doc comment onmap_tensor_data_cpusays so.Measurements
Peak
RssAnon, x86 Linux, Release,-DGGML_NATIVE=ON, 4 threads, one 8.6 s clip:The reduction tracks file size. Peak RSS is roughly unchanged — file pages still count toward RSS — but the anonymous share, which is what the OOM killer charges, drops by the size of the model.
On the device that motivated this (Galaxy S21 FE, Android 16), the 1.7B went from being reclaimed to running the voice-input panel with 0 lmkd kills.
Verification
Transcripts are byte-identical before and after on all three architectures. Dropping the staging copy makes load slightly faster (8.27 s → 7.38 s for the 1.7B on-device); the only cost is page-fault latency on the first run against a cold page cache — I measured 1.74 s then 1.22/1.23 s for parakeet, against 1.20 s unmapped.
Happy to split this per-arch, gate it behind an option, or drop the whisper/parakeet wiring if you'd rather land the helper alone first.