Skip to content
franknohPublic

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Linnet

Linnet is the checked source format for neural network architectures.

Think SafeTensors for model structure: the architecture lives in a typed .linnet file and the weights in SafeTensors. Linnet checks and loads the architecture without importing the model author's Python.

uv add "linnet-lang[torch]"                  # the compiler, its standard library and linnet.torch
linnet check model.linnet                    # shapes, dtypes, parameters; nothing runs
linnet inspect --parameters model.linnet     # every tensor a checkpoint must supply
from linnet.torch import load

model = load("model.linnet", generics={"H": 512, "Heads": 8}, weights="model.safetensors")

Documentation: linnet.franknoh.dev. Model registry: Nest. Try it in five minutes: the quickstart notebook on Colab.

Why Linnet

  • Architecture is source, weights are data. A .linnet file declares every parameter's shape and dtype. The PyTorch and ONNX loaders check each SafeTensors tensor against it before anything runs.
  • Git understands it. linnet fmt has one style, so a diff shows only model changes. linnet check in CI fails a change that breaks a shape downstream.
  • Run the model, not its repository. Loading reads declarative source and imports no code from the model's author, unlike trust_remote_code=True. It is not a sandbox, but the structure is known before anything runs.
  • Portable, not interpreted. Each backend gets its native path: generated PyTorch under torch.compile or CUDA graphs, XLA, ONNX Runtime and TensorRT.

Benchmarks

On one H100 in bf16:

Linnet Reference stacks
Llama 3.1 8B, decode one request 167 tok/s (XLA), 161 (CUDA graphs) vLLM 152, transformers compiled 110
BERT base, forward at batch 1 0.74 ms (CUDA graphs) transformers 3.60
Llama 3.1 8B, many requests at once 5600 tok/s vLLM 5449
Llama 3.1 8B, LoRA fine-tuning step 1.04 s TRL 1.83 s
Llama 3.1 8B, GRPO step 1.18 s TRL with vLLM 3.14 s

Every row, including where Linnet loses, is on the benchmarks page.

Targets

  • Load in PyTorch, JAX (XLA, jax.numpy, Flax NNX) and ONNX Runtime (CUDA, TensorRT).
  • Export StableHLO, ONNX, a Triton Inference Server model, a transformers checkpoint for vLLM, and GGUF for llama.cpp and Ollama.
  • Serve with linnet.serve: continuous batching and an OpenAI-compatible HTTP server.
  • Plan memory before running: linnet memory breaks down what a configuration needs, exact where the program decides it, and linnet fit finds the largest batch or context that fits a device.
  • Train in PyTorch or JAX: supervised fine-tuning, DPO and GRPO, with LoRA or fully sharded across GPUs. Or hand the model to transformers' Trainer and TRL.
  • Import from PyTorch, JAX, StableHLO and ONNX.
  • Run in ComfyUI with linnet-comfyui.

The compatibility matrix lists the entry points and limits of each target.

Nest

Nest is a registry of checked architectures and SafeTensors checkpoints (source). Its CI compiles every card, checks the published checkpoint against each parameter's shape and dtype, and exports the model to StableHLO, ONNX, PyTorch and JAX. linnet.nest.load loads a model by its Nest name, from any Hugging Face Hub repo with a card at its root, or from a directory. A transformers checkpoint of the Llama, Mistral, Qwen2, Qwen3, Phi-3 or GPT-2 families converts on the way: nest.load("Qwen/Qwen2.5-7B-Instruct").

Documentation

Contributing

CONTRIBUTING.md covers building, the checks CI runs and the conventions; SECURITY.md how to report a vulnerability. Changes are listed in CHANGELOG.md.

License

MIT; see LICENSE.

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages