
Build Kimi K3 from Scratch
Rebuild Moonshot's 2.8-trillion-parameter open model, one small module at a time.
Kimi Delta Attention, Gated MLA, Attention Residuals, Stable LatentMoE with 896 experts, MXFP4 quantization-aware training and a one-million-token context: every component of Kimi K3 built as a small, runnable PyTorch module, then checked against the released 2.8T checkpoint's own numbers and tensors.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Start Here3 capsules
What Kimi K3 is, what we will build instead of a 2.8T model, and the numbers that define the architecture — the map for the rest of the book.
01The model we are going to buildfree7 min02How to read this book🔒7 min03The shape of Kimi K32.8T🔒10 min01Foundations4 capsules
The pieces K3 shares with every transformer — tokenizer, embeddings, a standard attention block — built first so the K3-specific parts have something to differ from.
04The tokenizer and the embedding layer160K🔒8 min05A standard transformer block, as our baseline🔒7 min06What attention costs at one million tokensO(N²)🔒6 min07Linear attention: the state view🔒9 min02Kimi Delta Attention5 capsules
K3's workhorse layer, built from first principles: the delta rule, per-channel forgetting, short convolutions, and the full KDA module in PyTorch — 69 of the 93 layers.
08The delta rule from scratchKDAfree7 min09Forgetting, one channel at a time🔒10 min10Short convolutions, Swish and L2 normalization🔒9 min11The full KDA layer in PyTorchCODE🔒10 min12Position without RoPE, and chunkwise trainingNoPE🔒10 min03Gated MLA4 capsules
The 1-in-4 layers that keep exact global retrieval in the stack: DeepSeek-style latent KV compression, decoupled RoPE, and K3's output gate.
13Why keep any full attention at all3:1🔒7 min14Multi-head Latent Attention from scratchMLA🔒7 min15Decoupled RoPE: position in the MLA layersRoPE🔒7 min16The output gate🔒8 min04Attention Residuals2 capsules
K3's cross-depth pathway: instead of one residual stream accumulating everything, layers selectively retrieve from earlier depths through learned pseudo-queries.
17The residual stream and its limits🔒8 min18Attention Residuals from scratchAttnRes🔒14 min05Stable LatentMoE5 capsules
From a single MLP to 896 experts with 16 active: expert layers, the SiTU-GLU activation, latent routing, shared experts, and keeping the whole thing balanced.
19From one MLP to many expertsMoE🔒8 min20The expert MLP and SiTU-GLU🔒7 min21Routing: choosing 16 of 896TOP-16🔒7 min22Latent dispatch and shared experts3584🔒10 min23Load balancing and training stability🔒14 min06The Full Stack3 capsules
Everything assembled: the 93-layer layout in miniature, the parameter accounting that reproduces 2.8T and 104B, and a real training run of our mini-K3.
24Assembling mini-K3CODE🔒12 min25Counting to 2.8 trillionMATH🔒9 min26Training the miniatureRUN🔒8 min07MXFP4 and Quantization-Aware Training3 capsules
How a 2.8T model fits in 1.56 TB: microscaling formats from first principles, fake-quantization and the straight-through estimator, and QAT from SFT onward.
27MXFP4 and MXFP8 from first principlesMXFP4🔒9 min28Quantization-aware trainingQAT🔒9 min29What 1.56 TB on disk means🔒9 min08Long Context2 capsules
The million-token claim, examined: what the hybrid stack does to memory and speed at 1M tokens, and how to evaluate whether long context works.
30How the hybrid stack reaches one million tokens1M🔒9 min31Evaluating long context🔒8 min09Inference and Beyond4 capsules
Serving a 2.8T MoE in practice, the parts of K3 we did not rebuild, and what open weights make possible for you.
32Serving a 2.8T mixture of experts8×B300🔒13 min33Always-on reasoning and native vision🔒10 min34What open weights let you doOPEN🔒8 min35Where to go from here🔒6 minRatings & reviews
No ratings yet. Yours would be the first.