Omni-Models
Ask an early voice assistant a question and three separate programs answered in turn: one turned your speech into text, a language model wrote a reply, and a third read it aloud. It worked, but it felt like talking to a relay. An omni-model removes the relay.
What “omni” means
An omni-model handles several modalities in one system — text, audio, images, and video — instead of one model per type. The stronger form is any-to-any: any input can produce any output. Not just “understands images” but “sees a picture, hears a question about it, and answers aloud”.
The mechanism is shared representation. Each modality gets a modality encoder (and, for output, a decoder), and everything is projected into one common space the model reasons over. Sound is handled with speech tokens — small units of audio treated much like the text tokens from the tokenization chapter, so the same transformer can work with them.
Why one model beats three
- Latency. One hop instead of three. Speech-to-text-to-speech in a pipeline is inherently slower, and slower still when each stage waits for the last to finish.
- Prosody. A cascade throws away how you said it. Pitch, hesitation, sighing — everything that carries feeling — is flattened to text and lost. An omni-model hears it and can answer in kind.
- Interruption. Full-duplex means it listens while it speaks, so you can cut in like you would with a person. A cascade cannot easily do this, because its stages take turns.
Who ships them
The modern era of the form starts with GPT-4o (2024), which showed a single model understanding vision, audio and text and answering with speech. Gemini was built natively multimodal from the start. On the open side, Alibaba’s Qwen-Omni line (Qwen2.5-Omni, then Qwen3-Omni) reads text, images, audio and video and returns text or real-time speech; smaller efforts like Mini-Omni2 set out explicitly to reproduce a GPT-4o-like experience with open weights. That a frontier form arrived and was reproduced under a permissive licence within roughly a year is the same pattern the open-weights act described.
What it does not solve
Omni-models do not make any single modality better than a dedicated model — a specialist speech system can still win on sound quality, and a dedicated image model can still win on pictures. The gain is in the whole: latency, tone, and the ability to move between modalities without seams. And the hard problem is unchanged — how do you evaluate a system that sees, hears and speaks, when every benchmark so far measured one sense at a time? That question sits behind everything still to come: models that simulate worlds, and models that act in the physical one.