AvaritCall · Editorial team
Testing telephony STT with WER and critical field accuracy
Evaluate telephony STT with more than overall WER. Build a repeatable test using representative Turkish calls, matching codecs and separate checks for names, numbers and dates.
Test telephony STT with representative Turkish recordings from the actual call path. Measure WER and separately check names, order numbers and appointment dates. Use identical recordings, reference transcripts and scoring rules. Clean microphone recordings alone cannot establish telephone performance.
1. Build a dataset that represents Turkish phone conversations
The core dataset should contain calls resembling the actual workflow. For support, include product names, problem descriptions, addresses and order numbers. For reception, include names, dates, times and corrections. Label slow and fast speech, regional pronunciation, different devices, background conversation and interrupted sentences as separate slices. Include foreign brand names used within Turkish sentences.
If easy recordings dominate the pool, the overall average can conceal difficult calls. Write down the expected call distribution, then collect enough examples for each important slice. Evaluate short commands separately from long explanations. Controlled read recordings help probe particular sounds and numbers; natural telephone dialogue is also necessary. Split development and final evaluation pools by speaker and call session. Repeatedly using the final evaluation calls to tune settings undermines the value of that comparison.
2. Calculate WER against a common reference
WER = (S + D + I) / N. S counts substitutions, D deletions, I insertions and N reference words. Multiply by 100 to express a percentage. [1] NIST’s SCTK/sclite tool can align reference and system transcripts and count word errors. [2]
Synthetic calculation example: with N=100, S=5, D=3 and I=2, WER is 10%. For an actual evaluation, use human prepared and checked reference transcripts. Fix punctuation, capitalization, abbreviation and number formatting rules before scoring. Careless handling of Turkish I, ı, İ and i can distort the result. Retain raw and normalized output without changing scoring rules after seeing results. Calculate overall WER from summed errors and reference words, rather than averaging call percentages.
3. Measure fields that matter to the workflow
A filler word and an incorrect digit in a phone number can have similar influence on WER while affecting the task very differently. The field metrics below are proposed evaluation choices that expose this difference. Report the number correct and the total number evaluated with each metric. A percentage alone can create excessive confidence in a slice containing few examples.
| Field | Suggested metric | Evaluation rule |
|---|---|---|
| Names and entities | Correct fields / labeled fields | Use equivalences defined before scoring |
| Phone or order numbers | Exact sequences / labeled sequences | Preserve digit order and leading zeros |
| Dates and times | Correct interpretations / labeled expressions | Fix call date and time context |
| Incorrect critical confirmations | Tasks confirmed with incorrect information | Inspect the failure type per call |
Normalizing numbers to words or digits can make transcript comparison easier. It must not turn an incorrect number into an apparently correct one. Dates require similar care: scoring the expression tomorrow requires knowing when the call occurred. Keep the text recognition score separate from any later interpretation of the field.
4. Match the codec and sampling chain to real calls
The test audio’s codec, sample rate and capture point should match the live STT input. Changing a desktop recording’s file extension does not reproduce a telephone channel. Resampling a narrowband recording from 8 kHz to 16 kHz cannot restore missing frequencies; it calculates new samples from the existing signal. Torchaudio’s documentation describes bandlimited interpolation. [3]
In an illustrative appointment call, a user says Kadıköy’de yarın iki buçuk, then corrects the time with saat üç olsun. Record interim and final transcripts, and check whether the final date and time reach the appropriate fields. Compare the same call without network impairment and with controlled impairment. Determine whether the error arose in audio decoding, speech segmentation or processing the final correction. Follow file based evaluation with live call validation.
5. Checklist for a repeatable evaluation
- Record sample identifiers, call scenarios, speaker groups, codecs, sample rates and channel conditions in the dataset.
- Have a second listener review references. Agree on a marking and scoring rule for unintelligible sections before running the evaluation.
- Run every candidate on the same pool. Record model versions, language settings, dictionaries and every preprocessing step.
- Report scenario slices, critical fields and error examples alongside overall WER. Keep interim streaming output separate from final transcripts.
- Set acceptance thresholds according to the task before comparing results, then validate the outcome with new real call examples.
6. Frequently asked questions
Is the lowest WER always the best choice? No. A system that recognizes general conversation well can miss the names or numbers that appear frequently in your workflow. Complete the decision with field accuracy on representative tasks and the call experience.
Is a small test pool sufficient? It can expose the first obvious problems. Avoid generalizing until common and difficult slices each have enough examples. Show sample counts with the results, and inspect whether one speaker or one scenario dominates them.
How should a new STT version be compared? Run the locked evaluation pool again using identical rules. Listen to gains and regressions at call level. The overall average can conceal a critical field getting worse, so retain enough detail to understand the change.
- [1]KWS15 Keyword Search Evaluation Plan, section 6: Speech-to-Text Evaluation — National Institute of Standards and Technology (NIST), 2015-02-03 (accessed: 2026-10-01)
- [2]NIST SCTK: sclite scoring documentation — National Institute of Standards and Technology (NIST) (accessed: 2026-10-01)
- [3]Audio Resampling tutorial — PyTorch / Torchaudio (accessed: 2026-10-01)
