Voice technologyOct 1, 2026 · 5 min read

AvaritCall · Editorial team

Telephone echo and AEC: why your own voice comes back

Echo, feedback and double talk need different checks. Understand AEC reference audio and timing, then use speaker, microphone and overlapping speech trials to locate the problem.

Telephone echo and AEC: why your own voice comes back

Hearing your own voice after a delay can indicate echo; sustained squealing and speech cutting out during overlap call for different investigations. Acoustic echo cancellation, or AEC, aims to reduce the speaker sound that returns through a microphone. [1] First establish who hears what, then compare playback, microphone input and conversational timing for the same event.

1. Separate echo, feedback and double talk

Acoustic echo occurs when speaker sound enters the microphone and returns to the other party. [1] Feedback repeatedly amplifies microphone sound in a loop and can produce sustained squealing. [2] Double talk means speech from both sides overlaps in time. [3] Do not turn a report that “the audio comes back” into a configuration change before identifying which situation it describes.

Write concrete observations such as “the caller hears their last sentence again” or “the first syllable disappears only when both people speak”. A clean single-speaker trial cannot explain interruption behaviour. Repeat the same scenario with speakers exchanging roles to identify the affected direction. Preserve the sentence and listening method so the comparison remains useful.

2. Understand the AEC reference signal

The public Speex AEC interface uses microphone input together with the reference audio played through the speaker. [4] The basic objective is to separate desired near-end speech from returning far-end audio. Ask whether the reference represents the content actually played. An “outgoing audio” label from a different observation point is not enough to establish that relationship.

If an application produces a sentence but playback is cancelled, treating the entire generated sentence as heard can mislead an investigation. Record generation, playback start, stopping and microphone observations as separate events. This preserves the distinction between text appearing in a transcript and sound actually being heard in a room or through a telephone.

3. Check timing and the physical setup together

The timing relationship between reference and microphone signals matters; the Speex documentation identifies excessive delay and clipping as potential problems. [4] Do not rely solely on an “echo enabled” flag. Repeat a short sentence at a fixed level, then change one variable. Examine microphone position, playback level and device selection in separate trials so each result answers a specific question.

Comparing a headset trial with a speaker trial can offer evidence about the acoustic path, but cannot establish a software fault by itself. Describe the result as reproduced under that setup. Changing the connection and room at the same time makes attribution difficult. Document the environment and direction alongside the speech sample, including any changes that occurred between repetitions.

4. Evaluate near-end speech during double talk

ITU-T P.381 evaluates speech-level changes as well as echo during double talk. [3] Define success beyond reducing echo. When a caller interrupts with “no, Tuesday”, the first word must remain understandable. If returning audio becomes cleaner while nearby speech is suppressed, the conversation faces a different obstacle. Listen to the correction itself rather than judging only the background.

Three trials for the same conversation
TrialSymptom to listen forResult to record
One side speaksDelayed repetitionWhich party hears the echo
Both sides speakMissing first syllable or correctionIntelligibility for each speaker
Playback stopsAudio remaining after the stopAgreement between sound and event timing

WebRTC audio requirements recommend AEC or another form of echo control at endpoints. [5] This does not make every device behave identically. Link the evaluation to the audio setup being tested and reassess whether the result still applies after changing microphones or speakers.

5. Investigate a hypothetical reception call

A caller using speakerphone asks for a reservation time. During the greeting they interrupt with “wait, another day”, but the first word is unclear in the trial. Create examples with only the caller speaking, then only the greeting playing. Test the same correction during overlapping speech. The objective is to locate the conditions producing the symptom rather than classify every interruption as echo.

The next trial changes only playback level while keeping the previous setup. If the correction becomes understandable, record that result and assess remaining echo separately. Finally, compare the actual audible stop with its event timing. Microphone audio, speech detection and the output heard by the user should not collapse into a single success label. Each can explain a different part of the experience.

6. Use a repeatable AEC checklist

  • Identify the party hearing echo and the sentence that repeats.
  • Record speakers, microphone and headset arrangements for every trial.
  • Check how the reference relates to played content and its timing.
  • Evaluate single-speaker and overlapping speech separately.
  • Use short phrases with corrections, names and negation.
  • Make one change, then listen to echo and near-end intelligibility together.

Frequently asked questions

Can echo and squealing be fixed with the same setting? Identify the symptom first. Tracing a returning sentence and investigating a sustained feedback loop require different steps, even when both involve a microphone and speaker.

Does inaudible echo guarantee clean interruption handling? No. Evaluate first syllables and corrections during overlapping speech separately. Clean audio when only one person speaks cannot verify that behaviour by itself.

Sources
  1. [1]Audio for Distance Learning — Transmission Echo — Shure (accessed: 2026-10-01)
  2. [2]How to Control Feedback in a Sound System — Shure, 2013-01-25 (accessed: 2026-10-01)
  3. [3]ITU-T P.381 (03/2023): Technical requirements and test methods for analogue wired headsets or headphones and corresponding universal interface of terminals — ITU-T, 2023-03 (accessed: 2026-10-01)
  4. [4]Speex manual: Programming with Speex — Echo Cancellation — Speex / Xiph.Org (accessed: 2026-10-01)
  5. [5]RFC 7874: WebRTC Audio Codec and Processing Requirements — RFC Editor / IETF, 2016-05 (accessed: 2026-10-01)
Related solutions

See it on your own calls.

Set up in 5 minutes. $5 free on sign-up, pay as you go — no commitment.

Keep reading