Vizuara Books
LLM Production & Deployment
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · intermediate

LLM Production & Deployment

Take LLMs from notebook to production.

The full production stack for large language models: efficient serving with vLLM, text classification, clustering, prompt engineering, RAG, finetuning with LoRA and Unsloth, agents, and memory. Every stage is built hands-on the way real deployed systems are.

intermediatellmdeploymentproductionserving
44 capsules209 figures~10 hoursby Dr. Raj Dandekar

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00Foundations of Language AI5 capsules

Foundations of Language AI — 5 chapters.

01Why Deploying LLMs Is Hardconceptfree11 min02The History of Language AIconcept🔒12 min03Representation vs Generative Modelsconcept🔒13 min04Loading a Model with HuggingFacecode🔒13 min05Serving with vLLMcode🔒14 min
01Efficient Inference at Scale4 capsules

Efficient Inference at Scale — 4 chapters.

06What Makes vLLM Efficientdeep-dive🔒14 min07Throughput: vLLM vs HuggingFaceproject🔒13 min08Quantization for Local Inferencecode🔒13 min09Choosing a BERT-Based Modelintuition🔒14 min
02Text Classification with LLMs4 capsules

Text Classification with LLMs — 4 chapters.

10Classification with Representation Modelsconcept🔒14 min11Decoder & Encoder-Decoder Classificationconcept🔒14 min12Building a Text Classifiercode🔒11 min13Zero-Shot vs Finetuned Classificationintuition🔒15 min
03Clustering & Topic Modeling4 capsules

Clustering & Topic Modeling — 4 chapters.

14Text Clustering with Language AIconcept🔒12 min15Topic Modeling with Language AIconcept🔒12 min16Clustering & Topic Modeling Hands-Oncode🔒14 min17Project: Customer Support Clusteringproject🔒13 min
04Prompt Engineering in Production5 capsules

Prompt Engineering in Production — 5 chapters.

18Controlling Model Outputconcept🔒12 min19Basic Prompt Engineeringintuition🔒15 min20Advanced Prompting: CoT & ToTdeep-dive🔒12 min21Self-Consistency & Multiple Reasoning Pathscode🔒12 min22Evaluating Prompt Effectivenessproject🔒13 min
05RAG in Production6 capsules

RAG in Production — 6 chapters.

23The RAG Production Workflowconcept🔒14 min24Data Ingestion & Chunkingcode🔒14 min25Chunking Strategies Compareddeep-dive🔒15 min26Embedding & Retrievalcode🔒14 min27Generation & Groundingconcept🔒15 min28Building a RAG Web Appproject🔒14 min
06Finetuning LLMs9 capsules

Finetuning LLMs — 8 chapters.

29The Basics of LLM Finetuningconcept🔒17 min30Finetuning GPT-2 from Scratchcode🔒13 min31Tokenization & Padding for Finetuningcode🔒13 min32Finetuning with HuggingFace & Unslothcode🔒14 min33LoRA & Parameter-Efficient Finetuningmath🔒13 min34Soft-Prompt & Prefix Tuningcode🔒15 min35Multimodal Finetuning using JAXcode🔒15 min36Full vs LoRA Finetuning a Classifierproject🔒15 min37Finetuning Research: Subliminal Learning & RAFTdeep-dive🔒14 min
07Agents & Memory7 capsules

Agents & Memory — 7 chapters.

38Agents & the TAO Loopconcept🔒12 min39Building a Multi-Agent Frameworkcode🔒15 min40Orchestrating Agents with LangGraphcode🔒14 min41Agentic RAGdeep-dive🔒16 min42Project: Agents in Industryproject🔒15 min43Memory in LLM Systemsconcept🔒12 min44Project: Personalized Tutor with Mem0project🔒13 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book