© 2026

· SPEECH

You Cannot Hear a Fake. Why voice deepfake detection fails, and what replaces it

Deepfake detection is an arms race the detector loses. Stop asking “is this fake?” and ask “can this prove it is real?”

01 / THE SETUP

In January 2024 New Hampshire voters got a robocall in a synthetic copy of the US president’s voice telling them not to vote. Weeks later the FCC ruled such AI voices count as artificial voices under existing rules. The law moved fast. Detection did not.

02 / WHAT THE BENCHMARKS SAY

ASVspoof is the long-running spoofed-speech challenge. Its fifth edition, in 2024, used about 2,000 speakers in ordinary conditions, 32 attack algorithms and, for the first time, adversarial attacks.

The best systems aced their own test set. On older datasets and on In-the-Wild, real fakes of public figures from the internet, error rates rose sharply. The organisers’ verdict: detectors overfit to their datasets, and generalisation remains a holy grail.

03 / DIFFERENT, NOT HARDER

A 2024 study asked whether new fakes are harder or just different. Different: almost all of the drop on unseen attacks came from domain shift. A detector learns the fingerprints of generators it has seen. Each new one resets the game.

Phone codecs and compression make it worse, stripping the fine artefacts detectors need.

04 / FLIP THE QUESTION

Detection must prove a negative forever. Provenance proves a positive once: sign audio at capture and carry the signature with the file, as C2PA does, or watermark generated speech at the source, as Meta’s AudioSeal does.

Neither stops a determined attacker with an open model, but they change the default. Within a few years I expect unsigned audio to be treated like an unsigned email from your bank: not necessarily fake, just not evidence.

05 / THE CULTURAL SHIFT

For a century a recorded voice proved someone was there, speaking. That link is gone. We will trust chains of custody instead of timbre. Forged signatures did the same to writing, and we built notaries.