Vizuara Books
Build Kimi K3 from Scratch
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · advanced

Build Kimi K3 from Scratch

Rebuild Moonshot's 2.8-trillion-parameter open model, one small module at a time.

Kimi Delta Attention, Gated MLA, Attention Residuals, Stable LatentMoE with 896 experts, MXFP4 quantization-aware training and a one-million-token context: every component of Kimi K3 built as a small, runnable PyTorch module, then checked against the released 2.8T checkpoint's own numbers and tensors.

advancedllmmoemixture-of-expertskimilinear-attentionquantizationfrom-scratch
35 capsules148 figures~5 hoursby Dr. Raj Dandekar

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00Start Here3 capsules

What Kimi K3 is, what we will build instead of a 2.8T model, and the numbers that define the architecture — the map for the rest of the book.

01The model we are going to buildfree7 min02How to read this book🔒7 min03The shape of Kimi K32.8T🔒10 min
01Foundations4 capsules

The pieces K3 shares with every transformer — tokenizer, embeddings, a standard attention block — built first so the K3-specific parts have something to differ from.

04The tokenizer and the embedding layer160K🔒8 min05A standard transformer block, as our baseline🔒7 min06What attention costs at one million tokensO(N²)🔒6 min07Linear attention: the state view🔒9 min
02Kimi Delta Attention5 capsules

K3's workhorse layer, built from first principles: the delta rule, per-channel forgetting, short convolutions, and the full KDA module in PyTorch — 69 of the 93 layers.

08The delta rule from scratchKDAfree7 min09Forgetting, one channel at a time🔒10 min10Short convolutions, Swish and L2 normalization🔒9 min11The full KDA layer in PyTorchCODE🔒10 min12Position without RoPE, and chunkwise trainingNoPE🔒10 min
03Gated MLA4 capsules

The 1-in-4 layers that keep exact global retrieval in the stack: DeepSeek-style latent KV compression, decoupled RoPE, and K3's output gate.

13Why keep any full attention at all3:1🔒7 min14Multi-head Latent Attention from scratchMLA🔒7 min15Decoupled RoPE: position in the MLA layersRoPE🔒7 min16The output gate🔒8 min
04Attention Residuals2 capsules

K3's cross-depth pathway: instead of one residual stream accumulating everything, layers selectively retrieve from earlier depths through learned pseudo-queries.

17The residual stream and its limits🔒8 min18Attention Residuals from scratchAttnRes🔒14 min
05Stable LatentMoE5 capsules

From a single MLP to 896 experts with 16 active: expert layers, the SiTU-GLU activation, latent routing, shared experts, and keeping the whole thing balanced.

19From one MLP to many expertsMoE🔒8 min20The expert MLP and SiTU-GLU🔒7 min21Routing: choosing 16 of 896TOP-16🔒7 min22Latent dispatch and shared experts3584🔒10 min23Load balancing and training stability🔒14 min
06The Full Stack3 capsules

Everything assembled: the 93-layer layout in miniature, the parameter accounting that reproduces 2.8T and 104B, and a real training run of our mini-K3.

24Assembling mini-K3CODE🔒12 min25Counting to 2.8 trillionMATH🔒9 min26Training the miniatureRUN🔒8 min
07MXFP4 and Quantization-Aware Training3 capsules

How a 2.8T model fits in 1.56 TB: microscaling formats from first principles, fake-quantization and the straight-through estimator, and QAT from SFT onward.

27MXFP4 and MXFP8 from first principlesMXFP4🔒9 min28Quantization-aware trainingQAT🔒9 min29What 1.56 TB on disk means🔒9 min
08Long Context2 capsules

The million-token claim, examined: what the hybrid stack does to memory and speed at 1M tokens, and how to evaluate whether long context works.

30How the hybrid stack reaches one million tokens1M🔒9 min31Evaluating long context🔒8 min
09Inference and Beyond4 capsules

Serving a 2.8T MoE in practice, the parts of K3 we did not rebuild, and what open weights make possible for you.

32Serving a 2.8T mixture of experts8×B300🔒13 min33Always-on reasoning and native vision🔒10 min34What open weights let you doOPEN🔒8 min35Where to go from here🔒6 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book