A native Rust engine for running and training neural networks — hand-written CUDA kernels compiled at runtime via NVRTC, with no PyTorch, no libtorch, and no Python runtime.
Runs on any NVIDIA GPU from sm_80 (Ampere) up: kernels are JIT-compiled for the card,
with native NVFP4 / MXFP8 block-scale tensor-core paths on Blackwell (sm_120+) and
portable kernels everywhere else. Models load from single-file .syn bundles — quantized
to NVFP4, MXFP8, SQ1…SQ8 or kept dense, per layer group — or directly from GGUF, with
all 27 ggml block types executed as they are.
An alternative to the Python ML stack. Everything from the tensor API and CUDA kernels up to
full model ports, a tokenizer, an inference engine with paged KV-caches, and a training stack
is written in Rust: ~245k lines, 73 .cu kernel files, 2 000+ tests. Correctness is held to
bit-exact parity with PyTorch and NeMo reference implementations, per row rather than by a
global cosine similarity that hides local errors.
cargo build --release -p synaptix-cli
synaptix inspect model.syn # layout of a bundle
synaptix convert model.gguf model.syn # GGUF / safetensors → .syn (ggml blocks kept as is)
synaptix quantize model.syn model-sq4.syn --format sq4 # re-encode any weights (dense, NVFP4, ggml) into SQ4
synaptix run model.gguf "Explain NVFP4" --max-tokens 256 # .gguf loads directly
synaptix run model.syn "Explain NVFP4" --max-tokens 256 --quant sq4
synaptix chat model.syn --context 32768 # interactive, prefix-KV across turns
synaptix bench model.syn --n-tokens 128 # prefill / decode throughput
synaptix devices # compute capability, NVRTC target, block-scale MMA / TMA
synaptix run gemma4.syn "What is on the picture?" --image cat.png --chat
synaptix imagine sdxl.syn "a lighthouse at dusk" -o out.png # SDXL, FLUX.1, FLUX.2, Qwen-Image (+Edit), Qwen-Image 2.1
synaptix imagine flux2.syn "make it winter" --image photo.png # reference editing; --init-image/--strength for img2img
synaptix depth depth-anything-v2/ photo.png -o depth.png
synaptix video ltx.syn "a paper boat in the rain" -o clip.mp4 --gemma ./gemma-3-12b --seed 7
synaptix video ltx.syn "a singer on stage" --audio song.wav --upscaler up.syn # audio→video
synaptix h3 --model minimax-h3-fl2va.syn "waves at night" --first-frame shore.png
synaptix song "pop, female vocal" --lyrics-file lyrics.txt --seconds 90 --save-latent l.safetensors
synaptix song-decode l.safetensors --vae yue2-vae-legacy.syn -o song.wav
synaptix music "lofi piano, rain" -o track.wav --models ./syn_models --duration auto
synaptix speak voxcpm.syn "Hello there" -o out.wav --reference voice.wav
synaptix speak omnivoice.syn "Привет" --instruct "female, low pitch" --language ru
synaptix podcast vibevoice.syn "Speaker 1: hi\nSpeaker 2: hey" -o show.wav
synaptix transcribe whisper.syn talk.mp3 --format srt -o talk.srt # or gigaam.syn
synaptix diarize sortformer.syn meeting.wav --format rttm
synaptix embed bge-m3.syn "first text" "second text" -o vectors.json
synaptix rerank bge-reranker.syn "query" "doc one" "doc two" --top-k 1Native ports, each validated against its upstream reference:
| Domain | Models |
|---|---|
| LLM | Qwen3 (dense + MoE), Qwen3-Next hybrids (GatedDeltaNet + full attention, qwen3_5/3_6/3_8), Qwen4Exp (125B MoE: sparse-attention indexer, gated residuals, PLE n-grams, MTP head), Llama, Gemma-3, Gemma-4 26B A4B, Muse Glimmer 30B |
| Vision-language | Qwen3-VL tower (images and video, 3D M-RoPE), Gemma-4 vision tower, Muse Glimmer |
| Image | FLUX.1, FLUX.2 (dev, klein 4B / 9B), Qwen-Image 2.1 (text-to-image, editing by up to ten references, transparent RGBA), Qwen-Image and Qwen-Image-Edit (2509 / 2511: edit an image, or compose up to four references), SDXL (txt2img and img2img), Depth Anything V2 |
| Video | LTX-2.3 (22B), MiniMax-H3 (video with synchronized audio; image / video / audio references — Ref2VA) |
| Speech | Whisper, GigaAM (ASR), Sortformer (diarization) |
| Text-to-speech | VoxCPM, OmniVoice, VibeVoice (long-form, multi-speaker) |
| Music | YuE2 (song from style and lyrics through an editable ABC score; covers of a recording), SheetSage2 (recording → lead sheet in ABC: melody, chords, beats, key, sections), ACE-Step (generate, cover, edit, extend, extract, repaint) |
| Embeddings / rerank | BGE-M3, BGE-reranker-v2-m3 |
- NVFP4 (4-bit) and MXFP8 (8-bit) with block scaling, through
mma.synctensor-core instructions on Blackwell (sm_120). Quantization can be applied while packing a.synbundle, so the on-disk model is the deployed model. - SQ1…SQ8 — a portable block format (super-blocks of 256, bit-plane packing,
b + 0.625bits per weight) with a GPU encoder that searches the scale by MSE, bit-exact with its CPU reference. GGUF runs directly: llama/qwen2/qwen3/gemma3/gemma4 files load as they are, with all 27 ggml block types executed by the same portable kernels. - Any sm_80+ card. Kernels are JIT-compiled for the card's compute capability
(
synaptix devicesshows the target). Without Blackwell block-scale MMA, NVFP4 and MXFP8 weights run through in-register dequant GEMV / banded dequant + GEMM; a quantized bundle can also be transcoded on load (NVFP4 → SQ4, ggml → SQ, in RAM or a disk cache) by the loading policy.SYN_FORCE_ARCH=sm_80compiles everything for an older target to exercise that path on a newer card. - KV-cache in MXFP8 by default (per-layer: sliding-window layers stay unquantized), with block-table attention kernels that read the quantized cache directly.
- MoE offload — experts live in pinned host RAM and stream to the card on demand, with an arena allocator that returns VRAM to the driver on eviction instead of growing with the swap-in stream. A 125B MoE model runs on a 24 GB card.
- Partial block offload — N transformer blocks stay resident, the rest stream from the host, so a model larger than VRAM still runs (at a documented cost in tokens/s).
- Prefix-KV sessions — the KV of a conversation survives between turns for every architecture, including prompts that carry images, and can be parked in host RAM.
- Everything runs on a 7 GB card. Not "the small models": a 125B MoE chat model, a
27B hybrid, Gemma-4, FLUX.1 and FLUX.2, MiniMax-H3, LTX-2.3, ACE-Step with its 4B LM,
VibeVoice — each was measured with a ballast process holding all but 7 GB of VRAM.
Blocks that do not fit stream from pinned host RAM, or straight from the mmapped
.synbundle when RAM is short; embedding tables and heads stay on the host. The cost is bandwidth, and it is written down: FLUX.2 klein 1024² in 2–3 s, FLUX.1-dev 20 steps in 13 s, a 27B hybrid at 1–2 tok/s with 7 of 64 blocks resident.
Measured on an RTX 5090 Laptop (24 GB), 93 GB system RAM:
| Model | Prefill | Decode | Notes |
|---|---|---|---|
| Gemma-4 26B A4B | 10 100 tok/s @ 4k | 210 tok/s | CUDA-graph decode capturing the MoE, fused per-layer kernels |
| Qwen3.8-27B hybrid | 1 450 tok/s @ 3.3k | 47 tok/s | MTP speculative decode |
| Qwen3.8-Flash-Next 125B MoE | 1 650 tok/s @ 260k | 17–22 tok/s | 262k context on 24 GB; experts stream at ~39 GB/s |
Diffusion on the same card, 1024² (all weights resident):
| Model | Steps | Time | Peak VRAM |
|---|---|---|---|
| SDXL | 30 | 6.5 s | 8.8 GB |
| Qwen-Image 2.1, MXFP8 | 40 | 30 s | 13.6 GB |
| Qwen-Image 2.1, NVFP4 | 40 | 23 s | 9.3 GB |
| Qwen-Image 2.1 2048², MXFP8 | 40 | 172 s | 23.2 GB |
| Qwen-Image-Edit-2511, MXFP8 | 40 | 188 s | 17.8 GB |
| Qwen-Image-Edit-2511, NVFP4 | 40 | 155 s | 13.2 GB |
The 7 GB numbers quoted above are a different regime: most of the model is streamed, and the memory section says what stays resident.
Those are not starting points: Gemma-4 decode went 35 → 210 tok/s and prefill 957 → 10 100
tok/s over a week of kernel work (fused layer kernels, no memset on hot outputs, quantized
projections reading and writing BF16 directly, GQA-flash, windowed flash on sliding layers).
The write-ups live in the synthos repository under docs/.
Kernels are gated per-row against reference implementations. Performance is measured against
a maximally-tuned PyTorch baseline (torch.compile, FlashAttention, fp8), and the weaker
paths are documented rather than hidden — see LTX_GEMM_PARITY.md,
where bf16 GEMM lands at 0.82–1.16× of cuBLAS depending on shape: ahead on small and medium
M, behind on large-M tails.
| Crate | What it holds |
|---|---|
synaptix-core |
tensors, dtypes, devices, memory pools |
synaptix-kernels-cuda / -cpu |
73 .cu kernel files JIT-compiled via NVRTC; a CPU fallback |
synaptix-nn |
layers, attention, normalization, samplers |
synaptix-infer |
inference engine: paged KV, CUDA-graph capture, speculative decode |
synaptix-models |
the model ports listed above |
synaptix-bundle |
the .syn single-file format (mmap, zero-copy, quantize-on-pack) |
synaptix-tokenizer |
tokenizers, chat templates, tool-call parsers |
synaptix |
facades: llm, asr, tts, embedding, rerank, diarization, sampling |
synaptix-autograd / -train |
reverse-mode autograd, optimizers, checkpointing, eval |
synaptix-rag |
document parsing, chunking, retrieval helpers |
synaptix-cli |
the synaptix binary |
Honest scope note: inference is what is production-ready — it powers
synthos daily. The training side has a working
autograd, optimizers and checkpointing with tests, but the RLHF, distillation and self-play
modules are scaffolding, not finished trainers. Directories exist for models that are not
ported yet (they hold a stub lib.rs); the table above lists everything that actually runs.
Requires the CUDA toolkit (nvcc is read at build time to pin the CUDA version; the driver
itself is loaded dynamically at runtime through a
vendored cudarc patched to be safe under CUDA-graph capture).
cargo build --release -p synaptix-cli
cargo test --workspaceThe bit-exact test suite loads reference tensors that are not committed to this
repository — they are large and derived from upstream models. Regenerate them with the
scripts under scripts/reference/.
CUDA (primary) and CPU. Any sm_80+ card (Ampere, Ada, Hopper, Blackwell): NVRTC targets the
card's own compute capability. Native NVFP4 / MXFP8 block-scale mma.sync requires sm_120+
(Blackwell); on older cards the same weights run through dequantizing kernels, or are
transcoded to SQ on load. SQ and ggml formats use the same portable kernels on every card.
Young, single-author, and moving fast. The API is not stable; expect breaking changes.
This project is vibe-coded. Since spring 2026 I write all of my projects with Claude Code: I decide what to build and how it fits together, describe each task, and review, run and measure the result on my own hardware — the model writes the code, the tests and most of the documentation. Every number in this README was measured on my machine, and correctness is checked against reference implementations layer by layer rather than taken on the model's word.
synaptix is free and open source. If it is useful to you, you can support its development with a donation via PayPal.
Licensed under either of Apache-2.0 or MIT at your option.