Vizuara Books
Charlie and the Vision Room
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · intermediate · Charlie and the Intelligence Factory

Charlie and the Vision Room

Book II — Vision Transformers, from patches to pixels that generate themselves.

The second room. Convolution runs out of reach, so attention takes over: patches become tokens, ViT is coded from scratch, then DeiT, Swin, DETR, SAM, TimeSformer, NanoVLM, Flamingo and LLaVA, ending with autoencoders and diffusion. Every architecture is built by hand before it is used.

intermediatevisiontransformersvitmultimodaldetectiondiffusion
20 capsules119 figures~6 hoursby Dr. Raj Dandekar, Dr. Sreedath Panat

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00Into the Vision Room3 capsules

Why the machine that reads and the machine that sees ended up being the same machine.

01The Vision RoomSTARTfree17 min02Why convolution hits a wallCNN🔒16 min03The journey of a tokenTOKEN🔒18 min
01The Attention Bench4 capsules

Build attention four times, each time adding one thing you removed the time before.

04Simplified self-attentionATTN🔒17 min05Self-attention with trainable weightsQKV🔒19 min06Masked and multi-head attentionMHA🔒18 min07Counting the parametersMATH🔒16 min
02Building the Vision Transformer3 capsules

Cut the picture into patches and pretend they are words.

08Patches are tokensViT🔒17 min09Coding ViT from scratchCODE🔒16 min10DeiT and the distillation tokenDEIT🔒17 min
03The Hierarchy Room2 capsules

Attention over every patch is quadratic, which is fine for 196 patches and impossible for 50,000.

11Swin: windows that shiftSWIN🔒19 min12Coding Swin from scratchCODE🔒18 min
04Finding Things3 capsules

Classification is the easy task. Detection and segmentation are where the architecture has to change shape.

13DETR: detection as set predictionDETR🔒17 min14Coding DETR, and running itCODE🔒16 min15Segment AnythingSAM🔒16 min
05Time and Language3 capsules

Add a time axis, then add a text axis.

16TimeSformer and videoVIDEO🔒18 min17NanoVLM from scratchVLM🔒19 min18Flamingo and LLaVAMULTI🔒18 min
06Generating Pictures2 capsules

The last step is making the machine draw instead of describe.

19Autoencoders and VAEsVAE🔒18 min20Diffusion modelsDDPM🔒18 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book