
LLM Production & Deployment
Take LLMs from notebook to production.
The full production stack for large language models: efficient serving with vLLM, text classification, clustering, prompt engineering, RAG, finetuning with LoRA and Unsloth, agents, and memory. Every stage is built hands-on the way real deployed systems are.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Foundations of Language AI5 capsules
Foundations of Language AI — 5 chapters.
01Why Deploying LLMs Is Hardconceptfree11 min02The History of Language AIconcept🔒12 min03Representation vs Generative Modelsconcept🔒13 min04Loading a Model with HuggingFacecode🔒13 min05Serving with vLLMcode🔒14 min01Efficient Inference at Scale4 capsules
Efficient Inference at Scale — 4 chapters.
06What Makes vLLM Efficientdeep-dive🔒14 min07Throughput: vLLM vs HuggingFaceproject🔒13 min08Quantization for Local Inferencecode🔒13 min09Choosing a BERT-Based Modelintuition🔒14 min02Text Classification with LLMs4 capsules
Text Classification with LLMs — 4 chapters.
10Classification with Representation Modelsconcept🔒14 min11Decoder & Encoder-Decoder Classificationconcept🔒14 min12Building a Text Classifiercode🔒11 min13Zero-Shot vs Finetuned Classificationintuition🔒15 min03Clustering & Topic Modeling4 capsules
Clustering & Topic Modeling — 4 chapters.
14Text Clustering with Language AIconcept🔒12 min15Topic Modeling with Language AIconcept🔒12 min16Clustering & Topic Modeling Hands-Oncode🔒14 min17Project: Customer Support Clusteringproject🔒13 min04Prompt Engineering in Production5 capsules
Prompt Engineering in Production — 5 chapters.
18Controlling Model Outputconcept🔒12 min19Basic Prompt Engineeringintuition🔒15 min20Advanced Prompting: CoT & ToTdeep-dive🔒12 min21Self-Consistency & Multiple Reasoning Pathscode🔒12 min22Evaluating Prompt Effectivenessproject🔒13 min05RAG in Production6 capsules
RAG in Production — 6 chapters.
23The RAG Production Workflowconcept🔒14 min24Data Ingestion & Chunkingcode🔒14 min25Chunking Strategies Compareddeep-dive🔒15 min26Embedding & Retrievalcode🔒14 min27Generation & Groundingconcept🔒15 min28Building a RAG Web Appproject🔒14 min06Finetuning LLMs9 capsules
Finetuning LLMs — 8 chapters.
29The Basics of LLM Finetuningconcept🔒17 min30Finetuning GPT-2 from Scratchcode🔒13 min31Tokenization & Padding for Finetuningcode🔒13 min32Finetuning with HuggingFace & Unslothcode🔒14 min33LoRA & Parameter-Efficient Finetuningmath🔒13 min34Soft-Prompt & Prefix Tuningcode🔒15 min35Multimodal Finetuning using JAXcode🔒15 min36Full vs LoRA Finetuning a Classifierproject🔒15 min37Finetuning Research: Subliminal Learning & RAFTdeep-dive🔒14 min07Agents & Memory7 capsules
Agents & Memory — 7 chapters.
38Agents & the TAO Loopconcept🔒12 min39Building a Multi-Agent Frameworkcode🔒15 min40Orchestrating Agents with LangGraphcode🔒14 min41Agentic RAGdeep-dive🔒16 min42Project: Agents in Industryproject🔒15 min43Memory in LLM Systemsconcept🔒12 min44Project: Personalized Tutor with Mem0project🔒13 minRatings & reviews
No ratings yet. Yours would be the first.