© 2026

· SPEECH

The Voice Came Loose. The state of speech AI: cloning, codecs and super-resolution

Speech has become an editable medium, and a voice is now a parameter rather than a person.

01 / THREE SECONDS

In January 2023 Microsoft’s VALL-E cloned a voice from a three-second sample. It treated speech the way language models treat text: a sequence of tokens from a neural audio codec, predicted one after another. That framing took over the field.

By 2026 zero-shot cloning is a product feature. A few seconds of reference audio carries timbre, accent and rhythm, often across languages.

02 / THE CODEC IS THE MODEL

The quiet hero is the codec. SoundStream and EnCodec compressed audio into compact tokens, which made speech something a transformer could model.

Newer systems drop discrete tokens altogether. VoxCPM2, released this year, works in a continuous latent space. It encodes at 16 kHz and reconstructs at 48 kHz, so super-resolution falls out of the decoder for free. It was trained on over two million hours across 30 languages, and the weights are open.

03 / EDITING, NOT RECORDING

The bigger shift is control. Research codecs now try to separate timbre from prosody, so you can keep who is speaking and change how they say it. Systems like SpeechX fold noise removal, speaker extraction and word-level editing into one model.

None of this is lossless in the strict sense. Each edit is a fresh generation that the ear accepts as continuous. But perceptually it amounts to the same thing: audio is now as editable as text. You can fix a fluffed line after the fact and nobody hears the seam.

04 / REAL TIME

Conversation is the last stretch. Native speech-to-speech models like Kyutai’s Moshi answer in around 200 milliseconds and can listen while they talk. Most products still chain recognition, a language model and synthesis, trading speed for control. The gap between the two is now a design choice, not a research wall.

05 / AURA

Walter Benjamin wrote in 1935 that mechanical reproduction strips a work of art of its aura, its presence in one time and place. Recording did that to the voice once. Generation finishes the job. A voice no longer implies a throat.

I find that unsettling and beautiful. The voice becomes an instrument anyone can play, and the question of who is speaking moves from the sound to the context around it.