
Charlie and the Vision Room
Book II — Vision Transformers, from patches to pixels that generate themselves.
The second room. Convolution runs out of reach, so attention takes over: patches become tokens, ViT is coded from scratch, then DeiT, Swin, DETR, SAM, TimeSformer, NanoVLM, Flamingo and LLaVA, ending with autoencoders and diffusion. Every architecture is built by hand before it is used.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Into the Vision Room3 capsules
Why the machine that reads and the machine that sees ended up being the same machine.
01The Vision RoomSTARTfree17 min02Why convolution hits a wallCNN🔒16 min03The journey of a tokenTOKEN🔒18 min01The Attention Bench4 capsules
Build attention four times, each time adding one thing you removed the time before.
04Simplified self-attentionATTN🔒17 min05Self-attention with trainable weightsQKV🔒19 min06Masked and multi-head attentionMHA🔒18 min07Counting the parametersMATH🔒16 min02Building the Vision Transformer3 capsules
Cut the picture into patches and pretend they are words.
08Patches are tokensViT🔒17 min09Coding ViT from scratchCODE🔒16 min10DeiT and the distillation tokenDEIT🔒17 min03The Hierarchy Room2 capsules
Attention over every patch is quadratic, which is fine for 196 patches and impossible for 50,000.
11Swin: windows that shiftSWIN🔒19 min12Coding Swin from scratchCODE🔒18 min04Finding Things3 capsules
Classification is the easy task. Detection and segmentation are where the architecture has to change shape.
13DETR: detection as set predictionDETR🔒17 min14Coding DETR, and running itCODE🔒16 min15Segment AnythingSAM🔒16 min05Time and Language3 capsules
Add a time axis, then add a text axis.
16TimeSformer and videoVIDEO🔒18 min17NanoVLM from scratchVLM🔒19 min18Flamingo and LLaVAMULTI🔒18 min06Generating Pictures2 capsules
The last step is making the machine draw instead of describe.
19Autoencoders and VAEsVAE🔒18 min20Diffusion modelsDDPM🔒18 minRatings & reviews
No ratings yet. Yours would be the first.