AvaritCall · Editorial team
Streaming TTS: measuring time to first audio
First audio arrival, complete synthesis and audible playback happen at different times. Evaluate streaming TTS with a clear timing plan for text chunks, committed sentences and interruptions.
Streaming TTS lets playback use available audio before the whole response has been synthesized. Define the timing boundaries before measuring time to first audio: receiving audio after a TTS request and the user hearing a response are different events. Evaluate text chunking against sentence integrity and interruption behavior as well as speed.
1. Separate first audio, completion and audible playback
Microsoft documents first byte latency up to receiving the first audio chunk, and finish latency up to receiving all synthesized audio. [1] Define what TTFA means in your own evaluation. The first byte, the first decodable audio and the first audible sample can represent different boundaries. Avoid comparing another team’s TTFA with yours before checking its definition.
| Measurement | Start and end | Diagnostic purpose |
|---|---|---|
| First usable TTS audio | Request sent → usable audio received | Inspect generation and delivery startup |
| Playback startup | Request sent → first audible sound | Expose decoding and playback waits |
| Completion | Request sent → final audio received | Measure production of the complete response |
| Conversational response | User finishes → response becomes audible | Evaluate the complete conversational flow |
Do not confuse the audio’s playback duration with its generation time. A long reply can begin early while its final portion arrives later. Use boundaries measured against a common clock where possible, and record how timestamps from different devices are matched. Report slow trials and sample counts alongside an average.
2. Distinguish text chunks from audio chunks
Receiving chunked audio does not establish that an interface also accepts incremental text. Microsoft documents output streaming and input text streaming separately. [1] Check the chosen interface’s input behavior before designing the experiment. A received audio chunk need not coincide with a meaningful sentence boundary either.
Small text chunks create an opportunity to start sooner, but can split a word, number or stress pattern from its context. Larger chunks provide more context while waiting for additional text. Test short greetings, long explanations and spoken numbers separately rather than assuming one size works everywhere. Keep byte size separate from text length in the configuration and in the evaluation notes. Otherwise a change can affect several stages without making its cause clear.
3. Decide when a sentence is ready to be spoken
While a language model generates an answer, its final words may still be developing. Spoken information cannot be silently revised like text on a screen. Define a clear decision point for text ready to be voiced. Punctuation can provide a cue, but an incomplete date, unfinished number or unresolved negation can require another check.
For example, wait for the complete day and time expression before announcing an appointment. If an earlier start is useful, begin with a short verified sentence and deliver the detail in a separate sentence. Mark trials that voice incorrect or incomplete information, as well as trials with a fast first sound. Keep a traceable relationship between committed text, produced audio and the portion actually played. This makes a misleading response easier to locate.
4. Test an interrupted reception response
In an illustrative reception call, the assistant describes opening hours when the caller asks about another day. Stopping generation may be insufficient because already queued audio can still play. Record the start of caller speech, the interruption decision and the final audible sound from the old response as separate events.
Then check that the next response answers the new question. Listen for late audio from the old generation appearing within the new response. Repeat the trial with a short answer, a long explanation and two consecutive interruptions. The goal is early audible playback combined with correct sequencing when a response is interrupted. Add audible gaps and sentence integrity to the acceptance criteria. A fast first sound should not conceal an awkward recovery after the caller takes the turn.
5. Measurement and change checklist
- Record TTFA’s start and end events, clock basis and audio format with the test result.
- Measure first audio, playback startup and completion for the same response; do not combine times from different replies.
- Keep audio format, test sentence and network path fixed while changing text chunking.
- Listen to names, dates, numbers and negation. Mark audio that starts quickly but damages the intended meaning.
- Verify that old audio stops after interruption, late chunks remain associated with their response and the next answer plays in order.
- Observe first audio timing and playback queues as concurrency increases. Report the number of trials together with the results.
6. Frequently asked questions
Has the user heard the answer when the first byte arrives? Not necessarily. Audio must become decodable, enter playback and reach the receiver. Make the waiting time at those boundaries visible.
Should every word be sent to TTS immediately? A word boundary is not always a meaning boundary. Listen to the same sentence with different chunking and compare numbers, dates and stress.
Does lower TTFA make the whole call faster? It improves one stage. Measure the interval from the user finishing to audible response, and test interruption recovery separately, before judging the entire experience.
- [1]Lower speech synthesis latency using Speech SDK — Microsoft Learn, 2026-02-25 (accessed: 2026-10-01)

