
Charlie and the Language Room
Book I — Inference Engineering, from prefill to a million users.
The first room of the factory. Every trick that makes a large language model answer quickly and cheaply: prefill and decode, the KV cache and its costs, paged and sparse attention, FlashAttention, vLLM's scheduler, quantization, speculative decoding, parallelism, disaggregated serving, and the production layer that keeps a million users served.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00The Golden Ticket3 capsules
Where the factory is, why the door is worth queueing for, and the two halves of every answer a model gives.
01The Golden TicketSTARTfree7 min02Prefill and decodePHASES🔒8 min03The memory wallROOFLINE🔒7 min01The Cache Room4 capsules
The single data structure that makes generation tractable, and the same structure that eats all your memory.
04The good and evil of the KV cacheKV🔒7 min05Paged attentionVLLM🔒8 min06MQA, GQA and latent attentionHEADS🔒7 min07Sparse and sliding-window attentionSPARSE🔒7 min02The Attention Wing2 capsules
What happens when you refuse to store a cache at all, and what happens when you refuse to write the attention matrix to memory.
08State-space models and MambaSSM🔒7 min09FlashAttention 1, 2 and 3FLASH🔒8 min03The Engine Room3 capsules
Open vLLM and SGLang and look at the machinery that turns a pile of requests into a stream of tokens.
10The anatomy of a vLLM stepVLLM🔒17 min11Continuous batchingSCHED🔒15 min12Radix attention and prefix cachingSGLANG🔒17 min04The Squeezing Room4 capsules
Make the numbers smaller and guess the future. Two independent ways to buy speed.
13Quantization, from FP16 to FP4QUANT🔒17 min14GPTQ, AWQ and the outlier problemOUTLIER🔒19 min15Speculative decodingSPEC🔒16 min16Multi-token predictionMTP🔒17 min05The Scaling Room3 capsules
When one GPU is not enough, and when one machine should not be doing both jobs.
17Parallelism for inference5D🔒20 min18Disaggregated servingDISAGG🔒16 min19Distillation and finetuning for inferenceDISTILL🔒18 min06Out of the Factory3 capsules
The layer between a working engine and a service that a company can put its name on.
20Guardrails, guided decoding and evalsPROD🔒18 min21The million-user systemSCALE🔒17 min22Inference at the edgeEDGE🔒20 minRatings & reviews
No ratings yet. Yours would be the first.