AI Foundationspredict · compress · act
act XII

Beyond Text

For years, handling sound and pictures meant gluing three models together. An omni-model does it in one — and that changes how it feels to talk to.

54

Omni-Models

frontieras of 2026-09
before this →Image Models

Ask an early voice assistant a question and three separate programs answered in turn: one turned your speech into text, a language model wrote a reply, and a third read it aloud. It worked, but it felt like talking to a relay. An omni-model removes the relay.

THE OLD WAY — A CASCADEspeech → textlanguage modeltext → speechthree hops of delaytone and pauses lost at textTHE OMNI WAY — ONE MODELOMNI-MODEL — shared representationtext · audio · image · video intext · speech out, streamingone hop of delayprosody preservedcan be interruptedfull-duplex means listening and speaking at the same time, like a phone call rather than a walkie-talkie.
the cascade has three hops and loses the sound of your voice. The omni-model keeps everything in one system, so tone survives and answers can start before you finish.

What “omni” means

An omni-model handles several modalities in one system — text, audio, images, and video — instead of one model per type. The stronger form is any-to-any: any input can produce any output. Not just “understands images” but “sees a picture, hears a question about it, and answers aloud”.

The mechanism is shared representation. Each modality gets a modality encoder (and, for output, a decoder), and everything is projected into one common space the model reasons over. Sound is handled with speech tokens — small units of audio treated much like the text tokens from the tokenization chapter, so the same transformer can work with them.

Why one model beats three

Who ships them

2024 onward — and openly

The modern era of the form starts with GPT-4o (2024), which showed a single model understanding vision, audio and text and answering with speech. Gemini was built natively multimodal from the start. On the open side, Alibaba’s Qwen-Omni line (Qwen2.5-Omni, then Qwen3-Omni) reads text, images, audio and video and returns text or real-time speech; smaller efforts like Mini-Omni2 set out explicitly to reproduce a GPT-4o-like experience with open weights. That a frontier form arrived and was reproduced under a permissive licence within roughly a year is the same pattern the open-weights act described.

the honest noteThese model names will age within quarters — the page is stamped for that reason. The idea to keep is one shared representation instead of a relay of specialists, which is durable.
PIPELINEspecialists in a line, one hop each
OMNIone model, shared space, streaming

What it does not solve

Omni-models do not make any single modality better than a dedicated model — a specialist speech system can still win on sound quality, and a dedicated image model can still win on pictures. The gain is in the whole: latency, tone, and the ability to move between modalities without seams. And the hard problem is unchanged — how do you evaluate a system that sees, hears and speaks, when every benchmark so far measured one sense at a time? That question sits behind everything still to come: models that simulate worlds, and models that act in the physical one.

introduces →omni-modelany-to-anyfull-duplexspeech tokenmodality encoder
← previousImage Models