AvaritCall · Editorial team
VAD, end of turn and barge-in: when to listen and when to stop
A pause, a completed thought and an interruption are different events. Separate VAD, endpointing and barge-in decisions, then test conversational flow with pauses and corrections.
VAD finds speech activity in audio; deciding that a conversational turn is complete and stopping a playing response are separate tasks. Treating a pause as “the user has finished” can split an address or explanation. Treating every sound as an interruption can repeatedly break the conversation. Define the desired behaviour using example sentences first, then evaluate detection events, turn decisions and audible output separately.
1. Give the three decisions distinct names
Google Cloud STT provides speech activity begin and end events; events can also accompany audio yielding empty transcripts. [1] Such an event supplies an activity signal rather than proof that the user has completed their entire request. Endpointing chooses a boundary within audio. An end-of-turn decision asks whether the conversational context is ready for a response.
| Decision | Question | Separate verification |
|---|---|---|
| VAD | Is speech activity present? | Whether the audio actually contains speech |
| Endpointing / end of turn | Is this segment or turn complete? | Whether an ongoing sentence was split |
| Barge-in | Should the playing response stop? | Audible stopping and subsequent conversation behaviour |
This distinction also sharpens fault reports. Instead of “VAD is slow”, write “the start event occurred, but the response began while the user continued their sentence”. Detection and the decision based on detection then remain distinct, making it clear what the next trial needs to resolve.
2. Test speech onset with varied examples
Do not test only a clear, loud sentence. Prepare separate examples of a softly spoken correction, the first word after a breath, a short acknowledgment and background conversation. Write expected outcomes before testing: when should listening continue, when should playback stop, and when should the application ask again? Compare one sample under different decision rules so the result can be interpreted.
You need not choose between treating every microphone fluctuation as meaningful input and attending only to long sentences. A short “no” can matter greatly to the task. If the speaker or origin of an audible voice is uncertain, record that uncertainty too. Report clean-connection trials separately from trials with environmental sound.
3. Separate pauses from completed thoughts
Microsoft Speech documents a tradeoff: longer segmentation silence can accommodate pauses but delay results; shorter silence can split phrases. [2] This does not establish one universal duration. Someone saying “New … Town” may be searching for an address, whereas the information expected after “yes” may already be complete. The same silence can have different meanings in different tasks.
Deepgram’s documented endpointing feature uses VAD to detect pauses and emits speech_final. [3] Do not interpret that provider event as a universal guarantee of semantic completion. Include pauses after conjunctions, waits between list items and subsequent corrections in the test script. Count premature responses and unnecessary waiting as distinct failures, inspect the phrases causing them.
4. Compare interruption detection with audible stopping
RFC 6787 describes enabled Kill-On-Barge-In interrupting speech synthesis on speech onset or DTMF input. [4] When assessing a conversation, distinguish the request to stop from the moment audio actually stops. If a user says “no” but continues hearing the old response, inspecting speech onset alone does not explain the experience.
Test what happens after stopping. Does the old response resume, does the new correction change the answer, or is the user asked to repeat? Define suitable behaviour for each scenario. Evaluate recovery separately rather than combining unwanted interruptions and genuine corrections into a single success rate. The point is to preserve a useful conversation after the interruption, not merely produce a stop event.
5. A hypothetical address collection trial
A caller says “New … Town, Oak Street”; a response starts during the pause and the street name is missed. The team reads the same sentence fluently, with a natural searching pause, and with a short correction. It places speech onset, segment boundary decisions, response start and audible stopping on a timeline. The question is broader than which setting produces the fastest response.
Next, change only the end-of-turn decision while preserving the audio sample. If the street name survives but unnecessary waiting grows, record both outcomes. Then try “not street, avenue” while a response is playing. Acceptance means continuing with the corrected information as well as stopping promptly. This reveals behaviour that is fast but fragmented and behaviour that is coherent but slow, without compressing them into one judgment.
6. Use a conversational turn checklist
- Record the VAD event, turn decision and stopping outcome separately.
- Prepare examples with short acknowledgments, negation and quiet beginnings.
- Compare a pause inside a sentence with a genuinely completed response.
- Evaluate the first word of a correction arriving during playback.
- Change one decision rule and retest with the same examples.
- Report premature response, long waits, unwanted interruptions and recovery separately.
Frequently asked questions
Should a speech end event trigger an immediate reply? Evaluate the provider’s event definition together with task completion. Sound stopping and an explanation finishing are different observations.
What is the best silence threshold? No single number fits every call. Test values within the chosen interface’s limits using representative speech. Measure whether the conversation continues naturally at the appropriate moment, rather than declaring a universal timing rule.
- [1]Voice activity events and timeouts — Google Cloud, 2026-09-30 (accessed: 2026-10-01)
- [2]How to recognize speech — Change how silence is handled — Microsoft Learn, 2026-06-05 (accessed: 2026-10-01)
- [3]Endpointing — Deepgram (accessed: 2026-10-01)
- [4]RFC 6787: Media Resource Control Protocol Version 2 — Kill-On-Barge-In — RFC Editor / IETF, 2012-11 (accessed: 2026-10-01)
