
Build a NanoVLM from Scratch
Fuse vision and language from an empty file.
Build a tiny CLIP-style vision-language model by hand: encode images and captions into a shared embedding space and align them with contrastive learning. From patchified images and cosine similarity to the symmetric contrastive loss and zero-shot classification, every piece is coded and illustrated.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Vision-Language Foundations6 capsules
Vision-Language Foundations — 6 chapters.
01What Is a Vision-Language Model?conceptfree13 min02Why Multimodal — What VLMs Unlockintuition🔒13 min03One Idea: Vision and Language Are Both Vectorsintuition🔒12 min04From LLMs to VLMs: What Carries Overconcept🔒14 min05The Landscape: CLIP and Friendsconcept🔒15 min06What We Will Build: The NanoVLM Blueprintproject🔒15 min01Embeddings & Similarity6 capsules
Embeddings & Similarity — 6 chapters.
07Vectors as Meaningconcept🔒12 min08Text Embeddings from Scratchcode🔒15 min09Image Embeddings from Scratchcode🔒15 min10Cosine Similarity, Explainedmath🔒13 min11Dot Products, Norms, and Normalizationmath🔒13 min12The Shared Embedding Spaceintuition🔒12 min02Encoding Images into Vectors6 capsules
Encoding Images into Vectors — 6 chapters.
13Two Ways to Vectorize an Imageconcept🔒13 min14Patchify an Imagecode🔒15 min15Patch Embeddings and Positionscode🔒14 min16The CNN Route to an Image Vectorcode🔒15 min17Pooling to a Single Image Vectorcode🔒12 min18Matching Image and Text Dimensionscode🔒12 min03Contrastive Learning6 capsules
Contrastive Learning — 6 chapters.
19The Contrastive Ideaintuition🔒13 min20Positives, Negatives, and the Batchconcept🔒13 min21The Similarity Matrixmath🔒12 min22The CLIP Contrastive Lossmath🔒15 min23Temperature and Logit Scalingdeep-dive🔒14 min24Coding the Contrastive Losscode🔒12 min04Building the NanoVLM6 capsules
Building the NanoVLM — 6 chapters.
25NanoVLM Architecture Overviewconcept🔒14 min26The Image Encoder Modulecode🔒15 min27The Text Encoder Modulecode🔒14 min28The Projection Headscode🔒14 min29Synthetic Data for Trainingcode🔒14 min30Assembling the Full Forward Passcode🔒14 min05Training, Evaluation & Beyond6 capsules
Training, Evaluation & Beyond — 6 chapters.
31The Training Loopcode🔒16 min32Watching Alignment Emergeintuition🔒15 min33Zero-Shot Classificationproject🔒12 min34Image-Text Retrievalproject🔒14 min35Debugging and Scaling Tipsdeep-dive🔒16 min36From NanoVLM to Real VLMsconcept🔒12 minRatings & reviews
No ratings yet. Yours would be the first.