Linnet is the checked source format for neural network architectures.
Think SafeTensors for model structure: the architecture lives in a typed
.linnet file and the weights in SafeTensors. Linnet checks and loads the
architecture without importing the model author's Python.
uv add "linnet-lang[torch]" # the compiler, its standard library and linnet.torch
linnet check model.linnet # shapes, dtypes, parameters; nothing runs
linnet inspect --parameters model.linnet # every tensor a checkpoint must supplyfrom linnet.torch import load
model = load("model.linnet", generics={"H": 512, "Heads": 8}, weights="model.safetensors")Documentation: linnet.franknoh.dev. Model registry: Nest. Try it in five minutes: the quickstart notebook on Colab.
- Architecture is source, weights are data. A
.linnetfile declares every parameter's shape and dtype. The PyTorch and ONNX loaders check each SafeTensors tensor against it before anything runs. - Git understands it.
linnet fmthas one style, so a diff shows only model changes.linnet checkin CI fails a change that breaks a shape downstream. - Run the model, not its repository. Loading reads declarative source
and imports no code from the model's author, unlike
trust_remote_code=True. It is not a sandbox, but the structure is known before anything runs. - Portable, not interpreted. Each backend gets its native path:
generated PyTorch under
torch.compileor CUDA graphs, XLA, ONNX Runtime and TensorRT.
On one H100 in bf16:
| Linnet | Reference stacks | |
|---|---|---|
| Llama 3.1 8B, decode one request | 167 tok/s (XLA), 161 (CUDA graphs) | vLLM 152, transformers compiled 110 |
| BERT base, forward at batch 1 | 0.74 ms (CUDA graphs) | transformers 3.60 |
| Llama 3.1 8B, many requests at once | 5600 tok/s | vLLM 5449 |
| Llama 3.1 8B, LoRA fine-tuning step | 1.04 s | TRL 1.83 s |
| Llama 3.1 8B, GRPO step | 1.18 s | TRL with vLLM 3.14 s |
Every row, including where Linnet loses, is on the benchmarks page.
- Load in PyTorch, JAX (XLA,
jax.numpy, Flax NNX) and ONNX Runtime (CUDA, TensorRT). - Export StableHLO, ONNX, a Triton Inference Server model, a transformers checkpoint for vLLM, and GGUF for llama.cpp and Ollama.
- Serve with
linnet.serve: continuous batching and an OpenAI-compatible HTTP server. - Plan memory before running:
linnet memorybreaks down what a configuration needs, exact where the program decides it, andlinnet fitfinds the largest batch or context that fits a device. - Train in PyTorch or JAX: supervised fine-tuning, DPO and GRPO, with LoRA
or fully sharded across GPUs. Or hand the model to transformers'
Trainerand TRL. - Import from PyTorch, JAX, StableHLO and ONNX.
- Run in ComfyUI with linnet-comfyui.
The compatibility matrix lists the entry points and limits of each target.
Nest is a registry of checked architectures and
SafeTensors checkpoints (source). Its CI
compiles every card, checks the published checkpoint against each
parameter's shape and dtype, and exports the model to StableHLO, ONNX,
PyTorch and JAX. linnet.nest.load loads a model by its Nest name, from any
Hugging Face Hub repo with a card at its root, or from a directory. A
transformers checkpoint of the Llama, Mistral, Qwen2, Qwen3, Phi-3 or
GPT-2 families converts on the way: nest.load("Qwen/Qwen2.5-7B-Instruct").
- Installation and the quickstart
- Coming from PyTorch
- Language tour and specification
- LANGUAGE.md: the whole language in one file, for people and coding agents
- Command line
- PyTorch, JAX, ONNX, Nest, and integrations
- Training
- Benchmarks
CONTRIBUTING.md covers building, the checks CI runs and the conventions; SECURITY.md how to report a vulnerability. Changes are listed in CHANGELOG.md.
MIT; see LICENSE.