A reading voice that is not a punishment.
24 kHz
default output, full rate, not the telephony rate
Spell tags read a confirmation number back character by character, break tags make the structure audible, and a low temperature holds a steady delivery for someone listening all day. Read the whole document, every time.
Where the meter edits accessibility
Every accessible affordance is extra characters: the longer image description, the read-back of a confirmation number, the error explained rather than announced. On a meter each has a price, and the pressure is toward saying less.
audio
a screen-reader pass over a checkout summary
generated 2026-08-05
English
“Heading level two: your order summary. You have three items in the basket, and the total is sixty-four pounds and twenty pence including delivery. Below the summary there is a button labelled continue to payment, and after that a link back to the basket if you want to change anything first. Delivery is estimated for Thursday the fourteenth. To hear the individual items listed one by one, press the down arrow now.”
Mia · 21.2 s · accessibility · screen reader
What matters technically
- Spell tags read codes, references and confirmation numbers character by character, which is the difference between a usable read-back and a guess.
- A low temperature holds a steady delivery. That is the setting for a voice someone listens to all day.
- Break tags put a real beat between a heading and its content, so the structure is audible.
Notes
Is this a screen reader?
No. It is the layer a screen reader or your own application calls to turn text into audio. Navigation, semantics and focus stay in your product or in the assistive technology the person already uses.
What sample rate does the audio come back at?
24 kHz mono PCM by default, not the 8 kHz telephony band. Set output_sample_rate if your playback path wants something else.
A key and one stream to build this on, the same production API this page measures.
Get a key for this use case