
Build a Vision Transformer (ViT) from Scratch
Build the Vision Transformer by hand in PyTorch.
An image becomes a sequence of 16x16 patch tokens, and a plain transformer classifies it. Build every piece from scratch — patch embedding, class token, multi-head attention, encoder blocks, and the training loop.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Why Vision Transformers5 capsules
Why Vision Transformers — 5 chapters.
01Course Introductionconceptfree13 min02From CNNs to Transformersconcept🔒14 min03An Image is Worth 16x16 Wordsconcept🔒13 min04The ViT Architecture at a Glanceintuition🔒17 min05Setup and Tensor Conventionscode🔒14 min01Turning Images into Tokens6 capsules
Turning Images into Tokens — 6 chapters.
06Splitting an Image into Patchesintuition🔒15 min07Flattening and Linear Projectionmath🔒13 min08Patch Embedding with a Conv2d Trickcode🔒14 min09Coding the PatchEmbedding Classcode🔒12 min10The [CLS] Class Tokenconcept🔒13 min11Positional Embeddingsmath🔒12 min02Attention Inside the ViT6 capsules
Attention Inside the ViT — 6 chapters.
12Self-Attention Intuitionintuition🔒12 min13Queries, Keys, and Valuesmath🔒13 min14Coding Scaled Dot-Product Attentioncode🔒13 min15Multi-Head Attentionconcept🔒13 min16Coding Multi-Head Attentioncode🔒14 min17Why ViT Attention is Bidirectionaldeep-dive🔒13 min03The Transformer Encoder Block5 capsules
The Transformer Encoder Block — 5 chapters.
18Layer Normalizationmath🔒12 min19The MLP Feed-Forward Blockconcept🔒13 min20Residual Connectionsintuition🔒14 min21Coding One Encoder Blockcode🔒11 min22Stacking the Encodercode🔒13 min04Assembling the Full ViT5 capsules
Assembling the Full ViT — 5 chapters.
23The Classification Headconcept🔒11 min24Wiring the Full ViT Modelcode🔒12 min25Parameter Initializationcode🔒13 min26Counting Parameters and Computedeep-dive🔒12 min27End-to-End Forward Passcode🔒12 min05Training, Data, and Beyond6 capsules
Training, Data, and Beyond — 6 chapters.
28Why ViTs Are Data-Hungrydeep-dive🔒11 min29Loss, Optimizer, and the Training Loopcode🔒11 min30Training ViT on a Small Datasetproject🔒13 min31Visualizing Attention Mapsproject🔒13 min32Fine-Tuning Pretrained ViTsconcept🔒12 min33Putting It All Togetherproject🔒12 minRatings & reviews
No ratings yet. Yours would be the first.