ArticleslgStudy

science

Perceptual Evaluation of Speech Quality

Perceptual Evaluation of Speech Quality is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Perceptual Evaluation of Speech Quality rather than just read about it. In short: Perceptual Evaluation of Speech Quality (PESQ) is a family of standards comprising a test methodology for automated assessment of the speech quality as experienced by a user of a telephony system. It was standardized as Recommendation ITU-T P.862 in 2001.

Key takeaways

  • Perceptual Evaluation of Speech Quality belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Perceptual Evaluation of Speech Quality to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Perceptual Evaluation of Speech Quality from memory before moving on to harder problems.

Reference excerpt

Perceptual Evaluation of Speech Quality (PESQ) is a family of standards comprising a test methodology for automated assessment of the speech quality as experienced by a user of a telephony system. It was standardized as Recommendation ITU-T P.862 in 2001. PESQ is used for objective voice quality testing by phone manufacturers, network equipment vendors and telecom operators. Its usage requires a license. The first edition of PESQ's successor POLQA (Recommendation ITU-T P.863) entered into force in 2011.

Measurement scope PESQ was developed to model subjective tests commonly used in telecommunications (e.g., Recommendation ITU-T P.800) to assess the voice quality perceived by human beings. Consequently, it employs true voice samples as test signals. In order to characterize the listening quality as perceived by users, it is of paramount importance to load modern telecom equipment with speech-like signals. Many systems are optimized for speech and would respond in an unpredictable way to non-speech signals (e.g., tones, noise). Guidelines for proper applications of voice test samples are defined in the PESQ application guide contained in Recommendation ITU-T P.862.3.

Genealogy of related standards ITU-T's family of full reference objective voice quality measurements started in 1997 with Recommendation ITU-T P.861 (PSQM), which was superseded by ITU-T P.862 (PESQ) in 2001. P.862 was later complemented with Recommendations ITU-T P.862.1 (mapping of PESQ scores to a MOS scale), ITU-T P.862.2 (wideband measurements) and ITU-T P.862.3 (application guide). The first edition of ITU-T P.863 (POLQA) entered into force in 2011. An Application guide for Recommendation ITU-T P.863 was approved in 2019 and published as ITU-T P.863.1. In addition to the above listed full reference methods, the list of ITU-T's objective voice quality measurement standards also includes ITU-T P.563 (no-reference algorithm).

Testing typology Depending on the information that is made available to an algorithm, voice-quality test algorithms can be divided into two main categories:

A "full reference" (FR) algorithm has access to and makes use of the original reference signal for a comparison (i.e., a difference analysis). It can compare each sample of the reference signal (talker side) to each corresponding sample of the degraded signal (listener side). FR measurements deliver the highest accuracy and repeatability but can only be applied for dedicated tests in live networks (e.g., drive test tools for mobile network benchmarks). A "no reference" (NR) algorithm only uses the degraded signal for the quality estimation and has no information of the original reference signal. NR algorithms (e.g., Recommendation ITU-T P.563) are low-accuracy estimates only, as the originating voice characteristics (e.g., male or female talker, background noise, non-voice) of the source reference is completely unknown. A common variant of NR algorithms does not even analyze the decoded audio signal, but works on an analysis of the digital bit stream on an IP packet level. The measurement is consequently limited to a transport-stream analysis. PESQ is a full-reference algorithm and analyzes the speech signal sample-by-sample after a temporal alignment of corresponding excerpts of reference and test signal. PESQ can be applied to provide an end-to-end (E2E) quality assessment for a network, or characterize individual network components. PESQ results principally model mean opinion scores (MOS) that cover a scale from 1 (bad) to 5 (excellent). A mapping function to MOS-LQO is outlined in Recommendation ITU-T P.862.1.

See also Perceptual Objective Listening Quality Analysis (POLQA) Perceptual Evaluation of Video Quality (PEVQ) Perceptual Evaluation of Audio Quality (PEAQ) Hearing-Aid Speech Quality Index (HASQI)

References

Rix, Antony W.; Hollier, Michael P.; Hekstra, Andries P.; Beerends, John G. (2002-10-15). "Perceptual Evaluation of Speech Quality (PESQ) The New ITU Standard for End-to-End Speech Quality Assessment Part I--Time-Delay Compensation". Journal of the Audio Engineering Society. 50 (10): 755–764. Beerends, John G.; Hekstra, Andries P.; Rix, Antony W.; Hollier, Michael P. (2002-10-15). "Perceptual Evaluation of Speech Quality (PESQ) The New ITU Standard for End-to-End Speech Quality Assessment Part II: Psychoacoustic Model". Journal of the Audio Engineering Society. 50 (10): 765–778.

External links Application Note 1GA49: Psychoacoustic Audio Quality Measurements Using R&S UPV Audio Analyzer Application Note 1MA119: PESQ Measurement for GSM with R&SCMUgo Application Note 1MA136: PESQ Measurement for CDMA2000 with R&SCMUgo Application Note 1MA137: PESQ Measurement for WCDMA with R&SCMUgo Application Note 1MA149: VoIP Measurements for WiMAX Archived 2016-12-20 at the Wayback Machine

Worked examples

Example 1 — a first encounter with Perceptual Evaluation of Speech Quality

Start with the simplest possible case. Write down what Perceptual Evaluation of Speech Quality claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Perceptual Evaluation of Speech Quality before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Perceptual Evaluation of Speech Quality ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Perceptual Evaluation of Speech Quality

In research
Perceptual Evaluation of Speech Quality appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Perceptual Evaluation of Speech Quality in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Perceptual Evaluation of Speech Quality is common in secondary-school and first-year university syllabi. It links to neighbouring topics ITU-T P Series Recommendations, ITU-T recommendations, International standards, so understanding it makes those chapters shorter.
In everyday life
Look for Perceptual Evaluation of Speech Quality outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Perceptual Evaluation of Speech Quality in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Perceptual Evaluation of Speech Quality means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Perceptual Evaluation of Speech Quality out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Perceptual Evaluation of Speech Quality in simple terms?

Perceptual Evaluation of Speech Quality (PESQ) is a family of standards comprising a test methodology for automated assessment of the speech quality as experienced by a user of a telephony system. It was standardized as Recommendation ITU-T P.862 in 2001.

Why does Perceptual Evaluation of Speech Quality matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Perceptual Evaluation of Speech Quality?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Perceptual Evaluation of Speech Quality.

Tags

  • ITU-T P Series Recommendations
  • ITU-T recommendations
  • International standards
  • Speech codecs
  • Telecommunications

Keep exploring