Vizuara Books
Charlie and the Language Room
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · advanced · Charlie and the Intelligence Factory

Charlie and the Language Room

Book I — Inference Engineering, from prefill to a million users.

The first room of the factory. Every trick that makes a large language model answer quickly and cheaply: prefill and decode, the KV cache and its costs, paged and sparse attention, FlashAttention, vLLM's scheduler, quantization, speculative decoding, parallelism, disaggregated serving, and the production layer that keeps a million users served.

advancedinferenceservingllmkv-cachequantizationvllm
22 capsules103 figures~5 hoursby Dr. Raj Dandekar, Dr. Sreedath Panat

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00The Golden Ticket3 capsules

Where the factory is, why the door is worth queueing for, and the two halves of every answer a model gives.

01The Golden TicketSTARTfree7 min02Prefill and decodePHASES🔒8 min03The memory wallROOFLINE🔒7 min
01The Cache Room4 capsules

The single data structure that makes generation tractable, and the same structure that eats all your memory.

04The good and evil of the KV cacheKV🔒7 min05Paged attentionVLLM🔒8 min06MQA, GQA and latent attentionHEADS🔒7 min07Sparse and sliding-window attentionSPARSE🔒7 min
02The Attention Wing2 capsules

What happens when you refuse to store a cache at all, and what happens when you refuse to write the attention matrix to memory.

08State-space models and MambaSSM🔒7 min09FlashAttention 1, 2 and 3FLASH🔒8 min
03The Engine Room3 capsules

Open vLLM and SGLang and look at the machinery that turns a pile of requests into a stream of tokens.

10The anatomy of a vLLM stepVLLM🔒17 min11Continuous batchingSCHED🔒15 min12Radix attention and prefix cachingSGLANG🔒17 min
04The Squeezing Room4 capsules

Make the numbers smaller and guess the future. Two independent ways to buy speed.

13Quantization, from FP16 to FP4QUANT🔒17 min14GPTQ, AWQ and the outlier problemOUTLIER🔒19 min15Speculative decodingSPEC🔒16 min16Multi-token predictionMTP🔒17 min
05The Scaling Room3 capsules

When one GPU is not enough, and when one machine should not be doing both jobs.

17Parallelism for inference5D🔒20 min18Disaggregated servingDISAGG🔒16 min19Distillation and finetuning for inferenceDISTILL🔒18 min
06Out of the Factory3 capsules

The layer between a working engine and a service that a company can put its name on.

20Guardrails, guided decoding and evalsPROD🔒18 min21The million-user systemSCALE🔒17 min22Inference at the edgeEDGE🔒20 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book