Skip to content

Instant voice cloning

10 s

of reference, the whole enrollment

Zero-shot cloning builds a voice from a short reference clip at request time, no training job, no per-voice model. Ours takes about ten seconds of audio.

Instant vs trained cloning

Trained cloning fits a model to hours of a speaker’s audio, per voice, often per language. Zero-shot conditions a single model on a reference clip inside the request: send ten seconds, get that identity back. New voice, no pipeline.

table 01

The two cloning disciplines, side by side

Trained cloningZero-shot cloning
Enrollmenthours of audio, per voice~10 seconds, in the request
Waita training job, hours to daysnone, first request pays a cache fill
New languageoften a new training runsame reference, 23 languages
Cost per voicea per-voice fee or jobnothing marginal

The cross-language test

The hard case is a language the reference never spoke, where the identity has to survive without any of the original phonemes. That held up in blind listening panels. One reference carries one identity across 23 languages.

Every term on this page is measurable. Take a key and read the numbers off your own requests.

Get a key and measure it yourself