Vizuara Books
Vision-Language-Action & World Models for Robotics
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · advanced

Vision-Language-Action & World Models for Robotics

Build a self-driving VLA stack, from pixels to a real car.

A hands-on path through Vision-Language-Action models, world models, and reinforcement learning for autonomy: the Action Chunking Transformer, Isaac Lab, and code-as-policy. Every idea is built by hand and deployed onto a physical TurboPi car that drives itself.

advancedvlaworld-modelsroboticsautonomous
42 capsules215 figures~10 hoursby Dr. Raj Dandekar

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00Foundations: Perception, Language, Action6 capsules

Foundations: Perception, Language, Action — 6 chapters.

01Why Vision Alone Is Not Enoughintuitionfree15 min02The Perception-Reasoning-Control Stackconcept🔒16 min03Transformers and Self-Attention, Refreshedconcept🔒16 min04Vision Transformers: Images as Tokensconcept🔒15 min05Cross-Attention and the DETR Intuitionconcept🔒16 min06Setting Up Your Environmentcode🔒14 min
01Introduction to Vision-Language-Action5 capsules

Introduction to Vision-Language-Action — 5 chapters.

07What Is a Vision-Language-Action Model?concept🔒16 min08RT-2: A Vision-Language-Action Transformerdeep-dive🔒13 min09ACT: The Action Chunking Transformerdeep-dive🔒14 min10The VLA 3D Simulator Platformproject🔒14 min11Reading the RT-2 and ACT Papersdeep-dive🔒14 min
02Building the Action Chunking Transformer6 capsules

Building the Action Chunking Transformer — 6 chapters.

12ACT Architecture: A Conditional VAEdeep-dive🔒16 min13Action Chunking and Temporal Ensemblingmath🔒16 min14Collecting Demonstration Datacode🔒15 min15Training ACT in Simulationcode🔒14 min16The Nine ACT Experimentsproject🔒12 min17Evaluating and Debugging a VLA Policycode🔒13 min
03From Simulation to the TurboPi Car6 capsules

From Simulation to the TurboPi Car — 6 chapters.

18Self-Paced VLA Deep Divedeep-dive🔒14 min19Meet the TurboPi Platformconcept🔒13 min20TurboPi Commands and Controlcode🔒13 min21The Sim-to-Real Gapintuition🔒14 min22Deploying ACT on TurboPiproject🔒13 min23Closing the Loop: The First Autonomous Driveproject🔒14 min
04World Models5 capsules

World Models — 5 chapters.

24What Is a World Model?concept🔒12 min25Latent Dynamics and Imaginationmath🔒14 min26World Models for Drivingdeep-dive🔒12 min27World Model Experimentsproject🔒14 min28World Models vs Model-Free Policiesintuition🔒12 min
05Isaac Lab and Scalable Simulation5 capsules

Isaac Lab and Scalable Simulation — 5 chapters.

29Introduction to Isaac Labconcept🔒14 min30TurboPi in Isaac Labproject🔒12 min31Massively Parallel Simulationconcept🔒12 min32ACT Implementation in Isaac Labcode🔒14 min33The Isaac Lab Training Pipelineproject🔒14 min
06Reinforcement Learning for Self-Driving4 capsules

Reinforcement Learning for Self-Driving — 4 chapters.

34Why Reinforcement Learning for Driving?intuition🔒14 min35Reward Shaping for Autonomous Controlmath🔒11 min36Training a Driving Policy with RLcode🔒15 min37Combining Imitation and RLdeep-dive🔒13 min
07Code as Policy and Capstone5 capsules

Code as Policy and Capstone — 5 chapters.

38Code as Policy: The Ideaconcept🔒13 min39LLMs That Write Robot Programsdeep-dive🔒12 min40Building a Coding-Agent-as-Policyproject🔒12 min41The Full VLA Pipeline, End to Endproject🔒14 min42Capstone: Your Own Self-Driving Carproject🔒12 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book