简体中文 · Website · Installation · Server API · Build Log · Changelog · Public Alpha scope
0.1.0 Public Alpha · Qwen3.6-35B-A3B · text-only · CPU-native INT8
SparseFlow is a Qwen3.6-first research backend for tiered-memory sparse-expert inference. It keeps dense model state resident and gives routed MoE experts an explicit resident or SSD/cache storage policy.
The current release target is Qwen3.6-35B-A3B text-only Public Alpha.
Vision, MTP, shared streaming batching, and production serving are not part of
this release. A local OpenAI-compatible server is available for the frontend
integration workflow; see docs/api_contract.md.
The public project page is godiao.github.io/SparseFlow. It includes the system overview, measured release data, Route Trace concepts, and the build log behind this release.
67 GiB of model. About 3B active parameters per token. Native INT8 expert execution on a CPU, with routed experts kept on SSD when RAM is scarce.
SparseFlow is built around one practical question: how far can a sparse expert model run when the machine cannot keep every expert resident? The answer is not magic compression. It is an explicit storage hierarchy, a bounded cache, and a native kernel at the point where routed expert weights are actually used.
Qwen3.6-35B-A3B has 35B total parameters but activates only about 3B per token. Its routed experts account for roughly 60 GiB in the BF16 source checkpoint, while the dense resident core is about 6.97 GiB. SparseFlow turns that imbalance into a runtime boundary:
| Signal | Qwen3.6-35B-A3B |
|---|---|
| Total source checkpoint | 66.97 GiB |
| Routed experts, BF16 | 60.00 GiB |
| Routed experts, canonical INT8 container | 30.078 GiB |
| Dense resident core | about 6.97 GiB |
| Routed experts per layer | 256, top-8 selected |
| Transformer layers | 40 |
The model is still large. The useful idea is that the whole model does not need to occupy the same memory tier at the same time.
The host Transformers runtime continues to own attention, Gated DeltaNet, KV state, tokenizer handling, and generation orchestration. SparseFlow owns the model-independent expert boundary: locate a fused expert slice, lease it from the cache, read it from the INT8 container on a miss, and execute the routed expert with the native AVX-512 VNNI path.
| Capability | Release status |
|---|---|
| Qwen3.6 text-only generation | Public Alpha |
| Native INT8 resident hybrid | Stable baseline |
| Single-request INT8 SSD streaming | Stable low-memory path |
laptop-16gb preset |
Experimental and Doctor-gated |
| Local OpenAI-compatible Server | Available |
| SSE Chat and cancellation | Available |
| Runtime Doctor and memory guidance | Available |
| Route Trace | Available |
| Resident fixed-cohort grouped execution | Experimental opt-in |
| Shared streaming batching | Disabled after a measured NO-GO |
The first frontend release includes Overview, Chat, Route Trace, Models, and Runtime. Benchmark remains a separate workstream.
These views were captured from the real local SparseFlow Server. They show the current user-facing path: inspect the runtime, generate text, and follow the experts selected for a request.
The experiment-host matrix used an Intel Xeon Gold 6248R, 10 CPU threads, and NVMe storage. These are measured reference points, not universal performance guarantees:
| Path | TTFT | Decode | RSS | Expert reads |
|---|---|---|---|---|
| Native resident hybrid | 8.25 s | 2.4925 tok/s | 35.78 GiB | 0 |
| Native S1 streaming, 4 GiB cache | 20.21 s | 1.0981 tok/s | 9.62 GiB | 354.98 MiB/token |
| Native S1 streaming, 8 GiB cache | 18.01 s | 1.4157 tok/s | 13.72 GiB | 192.18 MiB/token |
| Model-cold S1, 4 GiB cache | 55.79 s median | 0.8632 tok/s median | <= 9.59 GiB | 355.0 MiB/token |
The resident and 4 GiB streaming native paths matched the frozen 60-question
quality matrix exactly for predictions and choice totals. Full methodology and
raw result links are in docs/results/qwen36_stage7_5_6_formal_20260716.md
and docs/results/qwen36_stage7_9_public_alpha_20260722.md.
The experimental laptop proof is intentionally reported separately: on a
15.85 GiB Windows laptop with AVX-512 VNNI, the laptop-16gb path completed a
real two-token generation at about 0.45 tok/s with roughly 1 GiB headroom after
closing other applications. It proves feasibility, not comfortable daily use.
The tested environments are Linux x86_64 and Windows 10 with Python 3.12,
PyTorch 2.9.x CPU, Transformers 5.x, Safetensors 0.8.x, and Accelerate 1.14.x.
The native W8A8 backend requires AVX-512 VNNI. CPU-only Python inspection and
planning commands do not require the runtime extras. See
docs/installation.md for the complete Windows/Linux
setup, storage budget, model preparation, Doctor checks, and server startup.
This is a source checkout release. The native C++ sources are compiled into a local cache on first native run; no model payload is copied to the Python package directory.
From the repository root:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[runtime]'
export SPARSEFLOW_NATIVE_CACHE="$PWD/.cache/native/int8_vnni"Set paths on a machine with enough disk space:
MODEL=/data/model/Qwen3.6-35B-A3B
INT8=/data/cache/qwen36-int8Check the model before reading payloads:
PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
--preset low-memory --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow inspect "$MODEL"
PYTHONPATH=src python -m sparseflow plan "$MODEL" --ram 96 --ctx 4096Prepare the versioned INT8 expert container. Conversion is shard/layer resumable and writes offline row sums required by the native path:
PYTHONPATH=src python -m sparseflow prepare-int8 "$MODEL" \
--output "$INT8" \
--report results/stage7_9_prepare_int8.json \
--json > results/stage7_9_prepare_int8.stdout.jsonRun the stable resident path:
PYTHONPATH=src python -m sparseflow run "$MODEL" \
--preset stable --int8-container "$INT8" \
--prompt "Explain sparse expert routing in one paragraph." \
--max-new-tokens 32 \
--output results/stage7_9_stable.jsonRun the stable single-request low-memory path:
PYTHONPATH=src python -m sparseflow run "$MODEL" \
--preset low-memory --int8-container "$INT8" \
--prompt "Explain sparse expert routing in one paragraph." \
--max-new-tokens 32 \
--output results/stage7_9_low_memory.jsonOn a roughly 16 GiB Windows laptop, use the explicit experimental profile. It requires AVX-512 VNNI and enough currently available RAM; Doctor must pass before a real runtime is started:
PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
--preset laptop-16gb --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow run "$MODEL" \
--preset laptop-16gb --int8-container "$INT8" \
--prompt "Explain sparse expert routing in one sentence." \
--max-new-tokens 2laptop-16gb is experimental and does not mean that every 16 GiB laptop is
supported. It uses a 256 MiB expert cache, 2048-token context, and a single
request. It is not the default preset and does not bypass the RAM admission
gate.
The fixed-cohort grouped path is explicit and experimental:
PYTHONPATH=src python -m sparseflow run "$MODEL" \
--preset experimental-batch --int8-container "$INT8" \
--prompt "Explain cache locality." \
--prompt "Explain MoE routing." \
--max-new-tokens 32The grouped preset requires equal encoded prompt lengths, as required by the fixed-cohort harness. Shared streaming batching is intentionally unavailable.
The tracked React/Vite frontend is under frontend/. It has no inference
backend of its own; start SparseFlow Server first, then use a second terminal:
$env:npm_config_cache = "$PWD\.cache\npm"
$env:PLAYWRIGHT_BROWSERS_PATH = "$PWD\.cache\playwright"
Set-Location frontend
npm ci
npm run devOpen the Vite URL shown in the terminal, normally http://127.0.0.1:5173.
The development proxy forwards /api/sparseflow to the local Python Server.
The production default is real Server mode. Fixture mode is only for explicit
UI development and must be labelled SIMULATED.
Chat supports explicit Qwen thinking control. Thinking is enabled by default;
use --disable-thinking when starting the Server to change the default, or
toggle Thinking in Chat Controls for one request. The request-level API field
is enable_thinking. The current Public Alpha streams model output as one text
channel; it does not expose a separate reasoning_content channel.
The release candidate checklist and final scope are in
docs/release_candidate.md.
The frozen 60-question HellaSwag/ARC/MMLU manifest can be scored with the same native hybrid dispatch used by the stable preset. The command stores choice scores and compact metadata, not generation logits:
PYTHONPATH=src python -m benchmarks.score_choices \
--model "$MODEL" \
--data benchmarks/manifests/quality_formal_v1.jsonl \
--backend int8-native --native-dispatch hybrid \
--int8-container "$INT8" --threads 10 \
--output results/stage7_9_quality_hybrid.jsonFor the low-memory quality path use --backend int8-native-streaming; keep
--native-dispatch hybrid and the default 4 GiB cache. The published 60-row
baseline is documented in
docs/results/qwen36_stage7_5_6_formal_20260716.md.
| Path | Status |
|---|---|
| INT8 native resident hybrid | Stable baseline |
| INT8 native single-request streaming S1 LRU | Stable low-memory |
| INT8 native laptop-16gb streaming | Experimental, host-dependent |
| INT8 native resident grouped fixed cohort | Experimental opt-in |
| Shared streaming batching/subcohort | Disabled, known limitation |
| Pure fused decode | Diagnostics only |
doctor validates model structure, safetensors headers, disk, CPU ISA,
container metadata, and optional native extension loading. The normal model
identity is a metadata plus payload-size digest; use --full-payload-hash only
when a complete 67 GiB payload hash is intentionally required.
PYTHONPATH=src python -m unittest discover -s tests -p 'test_*.py'The benchmark workstream owns formal quality manifests and raw result schemas. Large model payloads and caches stay outside Git; result JSON, reports, and manifests are committed.
- CPU native execution requires AVX-512 VNNI.
- The native extension is compiled from source on first use.
- Streaming is supported for one request at a time; prefetch is disabled in the Public Alpha preset.
- Grouped execution is a fixed equal-length cohort harness, not a dynamic scheduler.
- The host Transformers runtime still owns attention, Gated DeltaNet, KV state, tokenizer, sampling, and generation orchestration.
| Link | Purpose |
|---|---|
| Project website | Product overview and measured evidence |
| Installation guide | Linux and Windows setup, storage budget, and Doctor checks |
| Server API contract | Local OpenAI-compatible HTTP and SSE interface |
| Build Log | The engineering record behind the first release |
| Release scope | Public Alpha boundaries and acceptance criteria |
| Changelog | Versioned release notes |





