
Charlie and the Sound Room
Book III — Build voice agents from scratch, ear to mouth.
The third room. A real-time voice agent built component by component: sampling and spectrograms, voice activity detection, Whisper, streaming ASR, neural text-to-speech, the language model as an interlocutor, tool calling and memory, streaming buffers and barge-in, and the latency budget that decides whether the conversation feels alive.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Into the Sound Room3 capsules
A voice agent is not one model. It is a pipeline with a stopwatch running.
01The Sound RoomSTARTfree16 min02The cascaded pipelineARCH🔒17 min03What sound actually isDSP🔒16 min01Listening4 capsules
Getting text out of a moving audio stream, accurately and before the speaker has finished.
04Voice activity detectionVAD🔒18 min05How Whisper worksASR🔒16 min06Streaming ASR and faster-whisperSTREAM🔒19 min07Measuring transcription qualityWER🔒18 min02Speaking3 capsules
Turning a string back into a voice, fast enough that nobody notices the machine thinking.
08How text-to-speech worksTTS🔒18 min09Piper, Coqui and neural vocodersVOCODER🔒18 min10Prosody and streaming TTSPROSODY🔒22 min03The Brain4 capsules
The language model in the middle has to behave differently because it is being heard, not read.
11The LLM as an interlocutorLLM🔒17 min12Prompting for speechPROMPT🔒15 min13Tool calling for voice agentsTOOLS🔒19 min14Memory for voice agentsMEMORY🔒18 min04The Wire3 capsules
The plumbing that decides whether the agent feels quick or feels broken.
15Streaming, buffers and chunkingBUFFER🔒20 min16Barge-in and turn-takingBARGE🔒18 min17WebSockets and real-time transportNET🔒19 min05Shipping3 capsules
From a notebook to something a company would put on its phone line.
18Building Nimbus, end to endBUILD🔒19 min19The latency budgetLATENCY🔒19 min20Speech-to-speech, telephony and frameworksS2S🔒17 minRatings & reviews
No ratings yet. Yours would be the first.