AvaritCall · Editorial team
MOS, PESQ and POLQA: measuring telephone speech quality
MOS reflects listener ratings, while PESQ and POLQA use defined objective methods. Record the scale, version, reference audio and test context before comparing telephone speech quality.
Choose the question before measuring telephone speech quality. MOS summarizes listener judgments; PESQ and POLQA estimate quality using defined algorithms. A meaningful score needs a method, scale and test context. Add task accuracy and conversational timing when deciding whether clear sounding audio also preserves names and allows the conversation to progress comfortably.
1. Read subjective MOS with its test protocol
ITU-T P.800 defines subjective evaluation methods. In the common five category scale, 1 is bad, 2 poor, 3 fair, 4 good and 5 excellent; averaging ratings produces MOS. [1] State whether the experiment is a listening test or an interactive conversation test. Even with the same scale, those tasks describe different experiences.
Record listener count, language, playback equipment and audio level. Include the presentation order in the plan. Explain what is being rated: naturalness, audible impairment and intelligibility are different questions. A short practice session can help listeners interpret the scale consistently. Preserve more than one overall average when reporting the outcome. If a difficult recording produces divided opinions, that variation can matter to the decision as much as the mean.
2. Check the scope and current status of PESQ and POLQA
PESQ is associated with P.862 and POLQA with P.863. ITU states that the P.862 family was withdrawn on 5 January 2024 and points to P.863. [2] P.863 predicts listening quality; suitable reference and degraded signals matter. [3] Label older PESQ results with their historical method and version.
| Approach | Main input | Reporting detail |
|---|---|---|
| Subjective MOS | Listener ratings answering a defined question | Test type, scale and participant count |
| PESQ | Appropriate reference and processed speech | Legacy status, version and score mapping |
| POLQA | Appropriate reference and processed speech | P.863 version, operating mode and output definition |
| Task evaluation | Conversation and expected task outcome | Names, numbers, dates and response timing |
Do not give raw algorithm output, MOS mapped output and human MOS the same unqualified label. Read the measurement tool’s output definition and mapping. Check comparability before putting values from different versions or operating modes into one ranking. Describe the use case instead of adopting a universal threshold for acceptable speech quality.
3. Design a repeatable quality experiment
Pass the same reference speech through the audio path under evaluation and retain the receiver’s recording. The source and comparison recording must represent the same spoken content. Specify duration, sample rate, channel selection and trimming in the plan. Do not remove silence or damaged portions afterward to improve the score; define the evaluated portions in advance.
Change one factor in the initial trial: codec, network condition or processing step. Then check whether the difference repeats across speakers and sentences. Short digit strings and long explanations should each be represented deliberately. If objective results and listener judgments point in different directions, listen to the recordings together to identify the impairment affecting the decision. Keep that observation with the numeric result so the next comparison does not lose its explanation.
4. Combine quality and task results in a support call
In an illustrative support experiment, a caller states an order number and delivery day. Use the same sentences for two audio paths. The team first compares the recordings, then checks whether the number and day were recognized correctly. A segment that sounds fluent can still lose a critical digit. Speech quality and task accuracy are distinct findings in that case.
Also measure the interval between the caller finishing and the response starting. A listening quality measurement alone does not explain how that wait affects conversation. Keep separate rows for speech quality, critical field accuracy and conversational timing in the decision record. This makes strengths and weaknesses visible without forcing every observation into one score. The resulting decision can reflect the actual work the conversation must accomplish.
5. Measurement tool and report checklist
- State the standard, version, operating mode, score name and scale at the start of the report.
- Verify the reference source and matching spoken content, and record every transformation applied to the comparison audio.
- For listener tests, document language, equipment, level, rating question and participant count.
- Use the same test pool. Inspect distributions, sample counts and difficult recordings alongside the average.
- Evaluate critical word recognition and conversational delay separately from the quality score.
- Check the chosen implementation’s usage terms. OPTICOM publishes licensing information for PESQ and POLQA. [4] Access to a standard does not establish permission for every use of every software implementation.
6. Frequently asked questions
Can PESQ or POLQA score an arbitrary call recording? First check the reference and input conditions required by the implementation. Without an appropriate reference, another kind of measurement tool may be necessary.
Does a higher score imply more successful STT? They measure different properties. Run a separate recognition test covering important names, numbers and dates, then read its results together with the quality findings.
Are reports with the same MOS equivalent? Test type, participants, method, version and output scale must also provide compatible context. Numerical equality without those details is insufficient to establish an equivalent result.
- [1]ITU-T P.800: Methods for subjective determination of transmission quality — International Telecommunication Union (ITU), 1996-08-30 (accessed: 2026-10-01)
- [2]ITU-T P.862: recommendation status and withdrawal notice — International Telecommunication Union (ITU), 2024-01-09 (accessed: 2026-10-01)
- [3]ITU-T P.863 (03/2018): Perceptual objective listening quality prediction, summary — International Telecommunication Union (ITU), 2018-03-16 (accessed: 2026-10-01)
- [4]OEM technology licensing: PESQ and POLQA — OPTICOM (accessed: 2026-10-01)
