
5D Parallelism for Large Model Training
Scale a model across a cluster of GPUs.
Replicate the HuggingFace Ultra-Scale Playbook from scratch: data, tensor, sequence, context, and pipeline parallelism, plus ZeRO sharding and Mixture-of-Experts. Every axis is built by hand and profiled, then composed to distribute GPT-2 across a real GPU cluster.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00GPU Foundations6 capsules
GPU Foundations — 6 chapters.
01Why One GPU Is Not Enoughintuitionfree13 min02GPU Primer: Anatomy of an Acceleratorconcept🔒14 min03The Five Axes of Parallelismconcept🔒15 min04The Four Buckets of GPU Memoryconcept🔒14 min05Floating Point and Mixed Precisionmath🔒14 min06Memory Math for a 1B Llamamath🔒14 min01Data Parallelism7 capsules
Data Parallelism — 7 chapters.
07Activation Recomputationconcept🔒14 min08Gradient Accumulationconcept🔒13 min09Data Parallelism from Scratchcode🔒14 min10The Ring-All-Reduce Algorithmdeep-dive🔒13 min11Choosing the Batch Sizemath🔒14 min12Profiling with TensorBoardcode🔒13 min13Overlapping Communication and Computedeep-dive🔒15 min02ZeRO and Sharded Data Parallelism4 capsules
ZeRO and Sharded Data Parallelism — 4 chapters.
14The Redundancy Problem in Data Parallelismintuition🔒13 min15ZeRO-1: Sharding Optimizer Statesconcept🔒13 min16ZeRO-2 and ZeRO-3concept🔒15 min17Implementing ZeRO from Scratchcode🔒15 min03Tensor Parallelism5 capsules
Tensor Parallelism — 5 chapters.
18Why Split a Single Layerintuition🔒15 min19Column- and Row-Parallel Linear Layersmath🔒13 min20Tensor Parallelism for the Transformer Blockconcept🔒14 min21Tensor Parallelism from Scratchcode🔒15 min22TP vs ZeRO for Inferencedeep-dive🔒13 min04Sequence and Context Parallelism5 capsules
Sequence and Context Parallelism — 5 chapters.
23The Long-Sequence Memory Wallintuition🔒13 min24Sequence Parallelism Explainedconcept🔒14 min25Sequence Parallelism from Scratchcode🔒15 min26Context Parallelism and Ring Attentionconcept🔒14 min27Naive Context Parallelism from Scratchcode🔒15 min05Pipeline Parallelism5 capsules
Pipeline Parallelism — 5 chapters.
28Why Pipeline Parallelismintuition🔒14 min29The Pipeline Bubblemath🔒14 min30Advanced Pipeline Schedulesdeep-dive🔒14 min31Pipeline Parallelism from Scratchcode🔒16 min32Megatron 3D Parallelismdeep-dive🔒14 min06Expert Parallelism and MoE4 capsules
Expert Parallelism and MoE — 4 chapters.
33Mixture-of-Experts Basicsconcept🔒14 min34Expert Parallelism Explainedconcept🔒15 min35Load Balancing and Token Routingdeep-dive🔒16 min36Build an MoE SLM from Scratchproject🔒14 min07Putting It All Together4 capsules
Putting It All Together — 4 chapters.
37Composing 5D Parallelismconcept🔒14 min38Distributing GPT-2 Across a Clusterproject🔒14 min39Build an SLM with 5D Parallelismproject🔒15 min40Capstone Projects and Next Stepsproject🔒14 minRatings & reviews
No ratings yet. Yours would be the first.