Skip to content

How we measure TTS latency.

5 rules

behind every figure on this site

Every latency figure on this site comes from one harness with one set of rules, described here so you can hold us to it, or rerun it yourself.

The rules

  • Production API only, the endpoint customers ship on, not a staging box or a lab build.
  • Server time to first audio: request accepted to first playable byte, includes gateway, queue, and synthesis. Model-only time is never quoted.
  • Short conversational utterances with cloned voices, the harder case; long-form throughput is bounded below by it.
  • Both ends of the run, the median never ships without the range it came out of.
  • Conditions in full, every published figure states what it was measured under.

The run behind the figure

First audio byte in 146 ms over the open internet, 116 ms p50 first audio, server side warm.

Twelve draws give a low end, a median and a high end. They do not give a 95th percentile, so this site no longer prints one, an earlier figure did, and it was retired rather than re-derived from a sample too small to carry it.

What that looks like

In conversation people hand the turn back in about two tenths of a second. That window is the budget: first audio spends part of it, and your network and orchestration spend the rest.

What the clock includes

The single-stream figure this harness produces is server time to first audio. Here is exactly where that clock starts and stops:

The other clock

A client-measured figure reads higher than a server-side one because the network sits inside the clock. When this site publishes a comparison, every system in it is drawn from one vantage in one interleaved pass, and the chart says so.

The two clocks answer different questions. A comparison ranks systems against each other; the server-side figure sizes your own budget, and your network goes on top of it.

What we do not publish

Model-only clocks, staging builds, and a percentile our sample is too small to compute. A figure that cannot state its conditions does not ship. The console on the home page runs what you type on the production API, so the claim is checkable before you ever hold a key.

The harness is the public API. Take a key and rerun it with your script in it.

Get a key and rerun this yourself