Audio Models
The video chapter ended by noticing that new video models generate sound natively — the same network makes the picture and the speech. But that sound was never really modelled. It was attached. Speech was the first thing AI ever did. For most of its history it was done by a cascade: one system turned your voice into text, a language model wrote a reply, and a third system read it aloud. Every hop lost something — the timing, the tone, the fact that you had not finished your sentence.
An audio model is what you get when you stop treating sound as a detour through text and model it directly. The trick is the same one that worked for everything else.
Sound as tokens
The unlock is the neural audio codec — a model that learns to compress raw audio into a short sequence of discrete codes, the way an image model compresses a picture into a latent. Systems like SoundStream and EnCodec do this with residual vector quantisation: a coarse code captures the gist, then each further code fills in the detail the previous ones missed. A second of speech that was 16,000 numbers becomes a few dozen integers, and those integers are a speech token — the same kind of object a tokeniser produces for text.
That single move collapses two separate decades of engineering. Once sound is a sequence of tokens, you no longer need a speech-specific architecture: you need the language model you already have. The cascade’s three models become one, because speech recognition, translation, and speech synthesis all become next-token prediction in a shared space.
Three tasks, one recipe
Speech recognition — turning sound into text — was where scale first paid off loudly. OpenAI’s Whisper was trained on 680,000 hours of weakly supervised audio scraped from the internet, most of it not carefully transcribed. That volume bought robustness: accents, background noise, technical vocabulary and dozens of languages, from one model. The lesson was the same as everywhere else — enough messy data beats a little clean data.
Text-to-speech went the other way. VALL-E showed that a few seconds of a speaker’s voice, tokenised by a codec, could condition a language model to speak in that voice — zero-shot, with no fine-tuning per speaker. Naturalness stopped being the bottleneck; identity did, in both senses of the word.
Speech-to-speech is the version that feels different from a chat box. Systems like Moshi run a streaming codec (Mimi) inside a single model that listens and speaks at the same time — full-duplex, from the omni-model chapter — with roughly 200 ms of latency. That is fast enough to be interrupted, which every text interface in this guide cannot be. It is the first modality where the conversation includes overlapping with the other person.
Why music is harder
Music generation applies the same recipe to a harder target. MusicLM generated minutes of music from a text prompt by modelling a hierarchy of tokens — coarse tokens for long-range structure, fine tokens for fidelity — precisely because music has form that a second of speech does not. Text-to-music is real and shipped; the open problems are the ones a musician would name: coherent development over minutes, not just a plausible four bars.
What goes wrong
A voice is a credential, and cloning made it cheap. Three seconds of audio is enough to condition a synthesizer on a target speaker. That is the same capability that makes an audiobook in a dead relative’s voice possible, and it is the mechanism behind a fast- rising class of fraud: a phone call that sounds exactly like your child, your boss, or your bank. Voice is the modality people trust fastest and verify least. This is the security analogue of the deepfake problem, and it arrived earlier because audio needs far less signal than video.
Transcription fails quietly. A recognition model trained to always produce fluent text will produce fluent text for silence, music, or a language it does not know — it manufactures a sentence rather than returning nothing. The word error rate can look excellent on a benchmark and still hide this, because the benchmark contains no silence or noise. The standard accuracy metric does not count a confident invention as an error at all.
Evaluation is thin. Word error rate is a proxy for “did the words match”, not “was the meaning right” — punctuation, names, numbers, and negation all fail it. And like every benchmark in this guide, it measures what is easy to check: a well-resourced language read aloud, not a whispered aside in a noisy kitchen.
The frontier, late 2026
Closed. OpenAI’s voice stack, Google’s native-audio Gemini line, ElevenLabs for synthesis, and the audio side of every video model in the previous chapter.
Open. Whisper for recognition, MusicGen for music, Kyutai’s Moshi for real-time dialogue, and the audio arms of Qwen and MiniMax — the same pattern as images and video: the capability lands in open weights within a year or two.
The modality that talks back
Text is a request and a response. Sound is a conversation: it overlaps and it is interruptible; much of its meaning is in pitch and timing rather than words. That is why the cascade kept failing at the thing people actually wanted — not transcription, but talking.
Audio is also the point where the safety concerns of the next acts stop being hypothetical. A cloned voice is not a benchmark number; it is a tool for convincing a person that someone they love is in trouble. The modality AI handled first is the one where being wrong has the oldest, most human consequences.
Which is why the next chapter is a question rather than a technology. Once a model can hold a scene consistent, react to what you do, and keep its physics straight — pictures, sound and time all at once — what is it actually building?