A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
-
Updated
Oct 1, 2026 - Jupyter Notebook
A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
Build an LLM from scratch with transformers, attention, tokenization, dataset engineering, training, alignment, inference, multimodality, agents, and a runnable PyTorch mini language model.
Building a complete LLM from scratch. No libraries. No shortcuts. Just NumPy, math, and first principles.
Implementations, assignments, and experiment notes for Stanford CS336: Language Modeling from Scratch.
DeepSeek-style MoE + MLA from scratch in PyTorch, with a router-specialization probe (mutual information) and a dense/+MoE/+MLA ablation. Educational, nano-scale, measured.
Modular, step-by-step implementation of GPT & Modern LLMs from scratch in PyTorch. Featuring RoPE, RMSNorm, SwiGLU, GQA, KV-Cache, Unit Tests, and an Interactive Web Playground.
From-scratch PyTorch: frontier LLM techniques as of 2026-Q1 — the Muon optimizer and Multi-Token Prediction, plus a base BPE tokenizer. Self-contained, self-checking modules.
LLM Which Was Trained On A Consumer GPU
Modern decoder-only Transformer lab with GQA, RoPE, attention bridges, verified KV cache, and cloud experiments.
14 天读透一个 LLM 的全部实现:基于 MiniMind 的本地交互式学习站,逐行看 tensor 维度、可运行实验、小测
A large language model built from scratch using PyTorch. Currently, it is in the initial stage of the development process using PyTorch.🔧
A 0.51B bilingual small language model trained from scratch with SFT, GRPO, on-policy distillation, LoRA and VLM support.
A GPT-style language model trained from scratch on public-domain text using a decoder-only Transformer, built for LLM learning, experimentation, and research.
An industry-grade, production-ready Multi-Modal AGI Vision-Language LLM & Agentic Reasoning Framework built from scratch in PyTorch. Features ViT patch encoding, RoPE, SDPA Flash-Attention, KV-Cache, autonomous Plan-Act-Reflect agentic loop with tool dispatch & Google Colab T4 GPU support.
Transforming base GPT-2 (124M) into an instruction-following assistant on Databricks Dolly-15k in PyTorch.
Pretraining and fine-tuning a 124M GPT architecture from scratch on The Lord of the Rings corpus using PyTorch.
Adapting pretrained GPT-2 (124M) for 3-class financial news sentiment analysis by replacing the language modeling head with a sequence classification head in PyTorch.
A Transformer encoder built from first principles, implementing the core architecture from mathematical foundations to working PyTorch code, including tokenization, embeddings, positional encoding, self-attention, multi-head attention, LayerNorm, feed-forward networks, training, and evaluation.
A hands-on journey into building large language models from the ground up, exploring tokenization, embeddings, attention, Transformers, GPT, and the fundamentals behind modern LLMs.
To associate your repository with the llm-from-scratch topic, visit your repo's landing page and select "manage topics."