
Build a Data-Efficient Image Transformer (DeiT) from Scratch
Train a Vision Transformer on ImageNet-1k alone.
Build a Vision Transformer from patches to attention to a full encoder, then make it data-efficient with knowledge distillation and a dedicated distillation token. Every layer, loss, and training trick is coded by hand and illustrated.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Why DeiT?5 capsules
Why DeiT? — 5 chapters.
01Course Introductionconceptfree13 min02The Data Hunger of Vision Transformersintuition🔒15 min03What Makes DeiT Data-Efficientconcept🔒14 min04CNNs vs Transformers: The Inductive-Bias Tradeoffintuition🔒13 min05The Build Roadmapconcept🔒17 min01Images as Tokens5 capsules
Images as Tokens — 5 chapters.
06From Pixels to Patchesintuition🔒13 min07Patch Embedding with a Conv2dcode🔒13 min08The CLS Tokenconcept🔒14 min09Positional Embeddingsmath🔒12 min10Building the Embedding Blockcode🔒13 min02Attention from Scratch6 capsules
Attention from Scratch — 6 chapters.
11The Intuition of Self-Attentionintuition🔒14 min12Queries, Keys, and Valuesmath🔒15 min13Scaled Dot-Product Attention, Codedcode🔒13 min14Multi-Head Attentioncode🔒16 min15The MLP Block and GELUcode🔒13 min16LayerNorm and Residual Connectionsconcept🔒14 min03The ViT Backbone5 capsules
The ViT Backbone — 5 chapters.
17The Transformer Encoder Blockcode🔒13 min18Stacking the Encodercode🔒15 min19The Classification Headcode🔒15 min20Assembling the Full ViTproject🔒13 min21Counting Parameters and FLOPsdeep-dive🔒13 min04Knowledge Distillation6 capsules
Knowledge Distillation — 6 chapters.
22What Is Knowledge Distillation?concept🔒12 min23Soft Labels and Temperaturemath🔒13 min24The KL-Divergence Lossmath🔒15 min25Hard vs Soft Distillationconcept🔒13 min26Choosing the Teacherintuition🔒14 min27Combining the Lossescode🔒14 min05The Distillation Token5 capsules
The Distillation Token — 5 chapters.
28The Distillation-Token Ideaconcept🔒13 min29CLS Token vs Distillation Tokenintuition🔒12 min30Adding the Token to the Modelcode🔒16 min31The Two-Head Forward Passcode🔒15 min32Inference-Time Fusiondeep-dive🔒13 min06Training DeiT6 capsules
Training DeiT — 6 chapters.
33The Training Recipeconcept🔒14 min34Data Augmentation That Mattersintuition🔒13 min35Regularization: Stochastic Depth and Moreconcept🔒14 min36Writing the Training Loopcode🔒13 min37The Distillation Loss in Codecode🔒14 min38Evaluation and Fine-Tuningcode🔒16 min07Putting It Together4 capsules
Putting It Together — 4 chapters.
39The DeiT Model Variantsdeep-dive🔒14 min40Reproducing the Paper Resultsproject🔒14 min41Visualizing Attentiondeep-dive🔒13 min42From DeiT to the Frontierconcept🔒15 minRatings & reviews
No ratings yet. Yours would be the first.