AI Foundationspredict · compress · act

concepts → Task

Task

Text-to-speech

frontieras of 2026-093 connections

text in, a voice out — and with a few seconds of prompt, that specific voice

Where it sits

The prism has six jobs across and eight layers down. Its primary cell is no mode × L5. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Overlays: Evaluation & methodology.

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

nothing comes first — this is a starting pointText-to-speechnothing depends on it yet — a leaf in the reading order

What kind of thing it is — and what it is made of

It is part of: Audio model.

How it goes wrong — and what answers that

known failure modes
Voice cloninga few seconds of speech is enough to synthesise anyone — the failure mode where sound becomes a credential

Where to read it

The chapter that introduces it, and any chapter that uses it again.

60Audio Modelsact XIII · Beyond Text

Where it comes from

paperNeural Codec Language Models Are Zero-Shot Text to Speech SynthesizersChengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, Furu Wei · 2023

Every connection

All 3 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 1
is part ofAudio modelModel
Flow · 1
usesSpeech tokenQuantity
Failure · 1
fails byVoice cloningFailureMode

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 521 declared edges produce all 340 of these pages. Edit an edge and both endpoints change.