Skip to content

Repository files navigation

SparseFlow - tiered memory, sparse experts

0.1.0 Public Alpha · Qwen3.6-35B-A3B · text-only · CPU-native INT8

SparseFlow is a Qwen3.6-first research backend for tiered-memory sparse-expert inference. It keeps dense model state resident and gives routed MoE experts an explicit resident or SSD/cache storage policy.

The current release target is Qwen3.6-35B-A3B text-only Public Alpha. Vision, MTP, shared streaming batching, and production serving are not part of this release. A local OpenAI-compatible server is available for the frontend integration workflow; see docs/api_contract.md.

The public project page is godiao.github.io/SparseFlow. It includes the system overview, measured release data, Route Trace concepts, and the build log behind this release.

67 GiB of model. About 3B active parameters per token. Native INT8 expert execution on a CPU, with routed experts kept on SSD when RAM is scarce.

SparseFlow is built around one practical question: how far can a sparse expert model run when the machine cannot keep every expert resident? The answer is not magic compression. It is an explicit storage hierarchy, a bounded cache, and a native kernel at the point where routed expert weights are actually used.

Why SparseFlow

Qwen3.6-35B-A3B has 35B total parameters but activates only about 3B per token. Its routed experts account for roughly 60 GiB in the BF16 source checkpoint, while the dense resident core is about 6.97 GiB. SparseFlow turns that imbalance into a runtime boundary:

Signal Qwen3.6-35B-A3B
Total source checkpoint 66.97 GiB
Routed experts, BF16 60.00 GiB
Routed experts, canonical INT8 container 30.078 GiB
Dense resident core about 6.97 GiB
Routed experts per layer 256, top-8 selected
Transformer layers 40

The model is still large. The useful idea is that the whole model does not need to occupy the same memory tier at the same time.

The Core Idea

SparseFlow tiered-memory architecture

The host Transformers runtime continues to own attention, Gated DeltaNet, KV state, tokenizer handling, and generation orchestration. SparseFlow owns the model-independent expert boundary: locate a fused expert slice, lease it from the cache, read it from the INT8 container on a miss, and execute the routed expert with the native AVX-512 VNNI path.

What Is Real Today

Capability Release status
Qwen3.6 text-only generation Public Alpha
Native INT8 resident hybrid Stable baseline
Single-request INT8 SSD streaming Stable low-memory path
laptop-16gb preset Experimental and Doctor-gated
Local OpenAI-compatible Server Available
SSE Chat and cancellation Available
Runtime Doctor and memory guidance Available
Route Trace Available
Resident fixed-cohort grouped execution Experimental opt-in
Shared streaming batching Disabled after a measured NO-GO

The first frontend release includes Overview, Chat, Route Trace, Models, and Runtime. Benchmark remains a separate workstream.

Product Preview

These views were captured from the real local SparseFlow Server. They show the current user-facing path: inspect the runtime, generate text, and follow the experts selected for a request.

SparseFlow runtime overview

SparseFlow chat with route trace evidence

SparseFlow route trace timeline and physical reads

Runtime and route inspection details

SparseFlow logical route inspector

SparseFlow expert cache and output budget controls

SparseFlow Doctor RAM readiness view

Measured Release Data

The experiment-host matrix used an Intel Xeon Gold 6248R, 10 CPU threads, and NVMe storage. These are measured reference points, not universal performance guarantees:

Path TTFT Decode RSS Expert reads
Native resident hybrid 8.25 s 2.4925 tok/s 35.78 GiB 0
Native S1 streaming, 4 GiB cache 20.21 s 1.0981 tok/s 9.62 GiB 354.98 MiB/token
Native S1 streaming, 8 GiB cache 18.01 s 1.4157 tok/s 13.72 GiB 192.18 MiB/token
Model-cold S1, 4 GiB cache 55.79 s median 0.8632 tok/s median <= 9.59 GiB 355.0 MiB/token

The resident and 4 GiB streaming native paths matched the frozen 60-question quality matrix exactly for predictions and choice totals. Full methodology and raw result links are in docs/results/qwen36_stage7_5_6_formal_20260716.md and docs/results/qwen36_stage7_9_public_alpha_20260722.md.

The experimental laptop proof is intentionally reported separately: on a 15.85 GiB Windows laptop with AVX-512 VNNI, the laptop-16gb path completed a real two-token generation at about 0.45 tok/s with roughly 1 GiB headroom after closing other applications. It proves feasibility, not comfortable daily use.

Supported Environment

The tested environments are Linux x86_64 and Windows 10 with Python 3.12, PyTorch 2.9.x CPU, Transformers 5.x, Safetensors 0.8.x, and Accelerate 1.14.x. The native W8A8 backend requires AVX-512 VNNI. CPU-only Python inspection and planning commands do not require the runtime extras. See docs/installation.md for the complete Windows/Linux setup, storage budget, model preparation, Doctor checks, and server startup.

This is a source checkout release. The native C++ sources are compiled into a local cache on first native run; no model payload is copied to the Python package directory.

Quick Start

From the repository root:

python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[runtime]'
export SPARSEFLOW_NATIVE_CACHE="$PWD/.cache/native/int8_vnni"

Set paths on a machine with enough disk space:

MODEL=/data/model/Qwen3.6-35B-A3B
INT8=/data/cache/qwen36-int8

Check the model before reading payloads:

PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
  --preset low-memory --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow inspect "$MODEL"
PYTHONPATH=src python -m sparseflow plan "$MODEL" --ram 96 --ctx 4096

Prepare the versioned INT8 expert container. Conversion is shard/layer resumable and writes offline row sums required by the native path:

PYTHONPATH=src python -m sparseflow prepare-int8 "$MODEL" \
  --output "$INT8" \
  --report results/stage7_9_prepare_int8.json \
  --json > results/stage7_9_prepare_int8.stdout.json

Run the stable resident path:

PYTHONPATH=src python -m sparseflow run "$MODEL" \
  --preset stable --int8-container "$INT8" \
  --prompt "Explain sparse expert routing in one paragraph." \
  --max-new-tokens 32 \
  --output results/stage7_9_stable.json

Run the stable single-request low-memory path:

PYTHONPATH=src python -m sparseflow run "$MODEL" \
  --preset low-memory --int8-container "$INT8" \
  --prompt "Explain sparse expert routing in one paragraph." \
  --max-new-tokens 32 \
  --output results/stage7_9_low_memory.json

On a roughly 16 GiB Windows laptop, use the explicit experimental profile. It requires AVX-512 VNNI and enough currently available RAM; Doctor must pass before a real runtime is started:

PYTHONPATH=src python -m sparseflow doctor "$MODEL" \
  --preset laptop-16gb --int8-container "$INT8" --check-native
PYTHONPATH=src python -m sparseflow run "$MODEL" \
  --preset laptop-16gb --int8-container "$INT8" \
  --prompt "Explain sparse expert routing in one sentence." \
  --max-new-tokens 2

laptop-16gb is experimental and does not mean that every 16 GiB laptop is supported. It uses a 256 MiB expert cache, 2048-token context, and a single request. It is not the default preset and does not bypass the RAM admission gate.

The fixed-cohort grouped path is explicit and experimental:

PYTHONPATH=src python -m sparseflow run "$MODEL" \
  --preset experimental-batch --int8-container "$INT8" \
  --prompt "Explain cache locality." \
  --prompt "Explain MoE routing." \
  --max-new-tokens 32

The grouped preset requires equal encoded prompt lengths, as required by the fixed-cohort harness. Shared streaming batching is intentionally unavailable.

Local Frontend

The tracked React/Vite frontend is under frontend/. It has no inference backend of its own; start SparseFlow Server first, then use a second terminal:

$env:npm_config_cache = "$PWD\.cache\npm"
$env:PLAYWRIGHT_BROWSERS_PATH = "$PWD\.cache\playwright"
Set-Location frontend
npm ci
npm run dev

Open the Vite URL shown in the terminal, normally http://127.0.0.1:5173. The development proxy forwards /api/sparseflow to the local Python Server. The production default is real Server mode. Fixture mode is only for explicit UI development and must be labelled SIMULATED.

Chat supports explicit Qwen thinking control. Thinking is enabled by default; use --disable-thinking when starting the Server to change the default, or toggle Thinking in Chat Controls for one request. The request-level API field is enable_thinking. The current Public Alpha streams model output as one text channel; it does not expose a separate reasoning_content channel.

The release candidate checklist and final scope are in docs/release_candidate.md.

Quality Evaluation

The frozen 60-question HellaSwag/ARC/MMLU manifest can be scored with the same native hybrid dispatch used by the stable preset. The command stores choice scores and compact metadata, not generation logits:

PYTHONPATH=src python -m benchmarks.score_choices \
  --model "$MODEL" \
  --data benchmarks/manifests/quality_formal_v1.jsonl \
  --backend int8-native --native-dispatch hybrid \
  --int8-container "$INT8" --threads 10 \
  --output results/stage7_9_quality_hybrid.json

For the low-memory quality path use --backend int8-native-streaming; keep --native-dispatch hybrid and the default 4 GiB cache. The published 60-row baseline is documented in docs/results/qwen36_stage7_5_6_formal_20260716.md.

Public Alpha Path Status

Path Status
INT8 native resident hybrid Stable baseline
INT8 native single-request streaming S1 LRU Stable low-memory
INT8 native laptop-16gb streaming Experimental, host-dependent
INT8 native resident grouped fixed cohort Experimental opt-in
Shared streaming batching/subcohort Disabled, known limitation
Pure fused decode Diagnostics only

doctor validates model structure, safetensors headers, disk, CPU ISA, container metadata, and optional native extension loading. The normal model identity is a metadata plus payload-size digest; use --full-payload-hash only when a complete 67 GiB payload hash is intentionally required.

Testing

PYTHONPATH=src python -m unittest discover -s tests -p 'test_*.py'

The benchmark workstream owns formal quality manifests and raw result schemas. Large model payloads and caches stay outside Git; result JSON, reports, and manifests are committed.

Current Limitations

  • CPU native execution requires AVX-512 VNNI.
  • The native extension is compiled from source on first use.
  • Streaming is supported for one request at a time; prefetch is disabled in the Public Alpha preset.
  • Grouped execution is a fixed equal-length cohort harness, not a dynamic scheduler.
  • The host Transformers runtime still owns attention, Gated DeltaNet, KV state, tokenizer, sampling, and generation orchestration.

Project Links

Link Purpose
Project website Product overview and measured evidence
Installation guide Linux and Windows setup, storage budget, and Doctor checks
Server API contract Local OpenAI-compatible HTTP and SSE interface
Build Log The engineering record behind the first release
Release scope Public Alpha boundaries and acceptance criteria
Changelog Versioned release notes

About

Qwen3.6-first tiered-memory CPU inference for sparse expert models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages