AI Foundationspredict · compress · act
act XIII

Beyond Text

Sound is the modality with the oldest human interface and the longest pipeline. For years, speech was something you turned into text and back. The models that changed that made sound into tokens, exactly as text and images had been.

60

Audio Models

frontieras of 2026-09
before this →Omni-Models

The video chapter ended by noticing that new video models generate sound natively — the same network makes the picture and the speech. But that sound was never really modelled. It was attached. Speech was the first thing AI ever did. For most of its history it was done by a cascade: one system turned your voice into text, a language model wrote a reply, and a third system read it aloud. Every hop lost something — the timing, the tone, the fact that you had not finished your sentence.

An audio model is what you get when you stop treating sound as a detour through text and model it directly. The trick is the same one that worked for everything else.

THE WAVEFORM — 16,000 SAMPLES A SECONDcontinuous, enormous, and awkward for a transformer to eat directlyneural audio codeccompress to 1–6 kbpsdiscrete tokensa few dozen a second,not 16,000AFTER THE CODEC, THE SAME MACHINE READS BOTH“The cat sat …”text tokens▁▂▇▃▅▂▁▇▃speech tokensone transformernext-token predictionthe architecture does not knowit is listening. Predict the nexttoken, in text or in sound.this is the whole idea: get sound into the same token space as language, and the rest is already built.
a waveform is continuous, tens of thousands of samples a second. A neural codec compresses it into a short stream of discrete tokens — and after that, sound is a sequence problem, which is one we already know how to solve.

Sound as tokens

The unlock is the neural audio codec — a model that learns to compress raw audio into a short sequence of discrete codes, the way an image model compresses a picture into a latent. Systems like SoundStream and EnCodec do this with residual vector quantisation: a coarse code captures the gist, then each further code fills in the detail the previous ones missed. A second of speech that was 16,000 numbers becomes a few dozen integers, and those integers are a speech token — the same kind of object a tokeniser produces for text.

That single move collapses two separate decades of engineering. Once sound is a sequence of tokens, you no longer need a speech-specific architecture: you need the language model you already have. The cascade’s three models become one, because speech recognition, translation, and speech synthesis all become next-token prediction in a shared space.

Three tasks, one recipe

Speech recognition — turning sound into text — was where scale first paid off loudly. OpenAI’s Whisper was trained on 680,000 hours of weakly supervised audio scraped from the internet, most of it not carefully transcribed. That volume bought robustness: accents, background noise, technical vocabulary and dozens of languages, from one model. The lesson was the same as everywhere else — enough messy data beats a little clean data.

Text-to-speech went the other way. VALL-E showed that a few seconds of a speaker’s voice, tokenised by a codec, could condition a language model to speak in that voice — zero-shot, with no fine-tuning per speaker. Naturalness stopped being the bottleneck; identity did, in both senses of the word.

Speech-to-speech is the version that feels different from a chat box. Systems like Moshi run a streaming codec (Mimi) inside a single model that listens and speaks at the same time — full-duplex, from the omni-model chapter — with roughly 200 ms of latency. That is fast enough to be interrupted, which every text interface in this guide cannot be. It is the first modality where the conversation includes overlapping with the other person.

Why music is harder

structure over minutes, not milliseconds

Music generation applies the same recipe to a harder target. MusicLM generated minutes of music from a text prompt by modelling a hierarchy of tokens — coarse tokens for long-range structure, fine tokens for fidelity — precisely because music has form that a second of speech does not. Text-to-music is real and shipped; the open problems are the ones a musician would name: coherent development over minutes, not just a plausible four bars.

the honest limitSpecialist speech systems still beat general models on raw sound quality, and singing is not solved by the same trick as talking. “Omni” means one model can do it, not that one model does everything best.
SPEECHmilliseconds — tone, timing, interruption
MUSICminutes — form and development
EVERYTHINGone token space for both

What goes wrong

A voice is a credential, and cloning made it cheap. Three seconds of audio is enough to condition a synthesizer on a target speaker. That is the same capability that makes an audiobook in a dead relative’s voice possible, and it is the mechanism behind a fast- rising class of fraud: a phone call that sounds exactly like your child, your boss, or your bank. Voice is the modality people trust fastest and verify least. This is the security analogue of the deepfake problem, and it arrived earlier because audio needs far less signal than video.

Transcription fails quietly. A recognition model trained to always produce fluent text will produce fluent text for silence, music, or a language it does not know — it manufactures a sentence rather than returning nothing. The word error rate can look excellent on a benchmark and still hide this, because the benchmark contains no silence or noise. The standard accuracy metric does not count a confident invention as an error at all.

Evaluation is thin. Word error rate is a proxy for “did the words match”, not “was the meaning right” — punctuation, names, numbers, and negation all fail it. And like every benchmark in this guide, it measures what is easy to check: a well-resourced language read aloud, not a whispered aside in a noisy kitchen.

The frontier, late 2026

a token space, now shared

Closed. OpenAI’s voice stack, Google’s native-audio Gemini line, ElevenLabs for synthesis, and the audio side of every video model in the previous chapter.

Open. Whisper for recognition, MusicGen for music, Kyutai’s Moshi for real-time dialogue, and the audio arms of Qwen and MiniMax — the same pattern as images and video: the capability lands in open weights within a year or two.

what to watchProvenance and consent: watermarking and detection for synthetic speech, and whether “a voice is a credential” is treated as a security problem rather than a novelty. The technology is not the part that is undecided.
WAVEcontinuous pressure — unmanageable
TOKENScodec makes sound a sequence
SAME MODELlisten, speak, interrupt

The modality that talks back

Text is a request and a response. Sound is a conversation: it overlaps and it is interruptible; much of its meaning is in pitch and timing rather than words. That is why the cascade kept failing at the thing people actually wanted — not transcription, but talking.

Audio is also the point where the safety concerns of the next acts stop being hypothetical. A cloned voice is not a benchmark number; it is a tool for convincing a person that someone they love is in trouble. The modality AI handled first is the one where being wrong has the oldest, most human consequences.

Which is why the next chapter is a question rather than a technology. Once a model can hold a scene consistent, react to what you do, and keep its physics straight — pictures, sound and time all at once — what is it actually building?

introduces →audio modelspeech recognitiontext-to-speechspeech-to-speechneural audio codecmusic generationvoice cloning
← previousVideo Modelsnext →World Models