ArticleslgStudy

physics

MUSHRA

MUSHRA is a physics topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand MUSHRA rather than just read about it. In short: MUSHRA stands for Multiple Stimuli with Hidden Reference and Anchor and is a methodology for conducting a codec listening test to evaluate the perceived quality of the output from lossy audio compression algorithms. It is defined by ITU-R recommendation BS.1534-3.

Key takeaways

  • MUSHRA belongs to physics; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect MUSHRA to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of MUSHRA from memory before moving on to harder problems.

Reference excerpt

MUSHRA stands for Multiple Stimuli with Hidden Reference and Anchor and is a methodology for conducting a codec listening test to evaluate the perceived quality of the output from lossy audio compression algorithms. It is defined by ITU-R recommendation BS.1534-3. The MUSHRA methodology is recommended for assessing "intermediate audio quality". For very small or sensitive audio impairments, Recommendation ITU-R BS.1116-3 (ABC/HR) is recommended instead. MUSHRA can be used to test audio codecs across a broad spectrum of use cases: music and film consumption, speech for e.g. podcasts and radio, online streaming (in which trade-offs between quality and efficiency of size and computation are paramount), modern digital telephony, and VOIP applications (which require quasi-real-time, low-bitrate encoding that remains intelligible). Professional, "audiophile", and "prosumer" uses are typically better suited to alternative tests, like the aforementioned ABC/HR, with a base assumption of high-quality, high-resolution audio wherein there will be minimal detectable differences between reference material and the codec output. The main advantage over the mean opinion score (MOS) methodology (which serves a similar purpose) is that MUSHRA requires fewer participants to obtain statistically significant results. This is because all codecs are presented at the same time, to the same participants, such that a paired t-test or repeated measures analysis of variance can be used for statistical analysis. Furthermore, the 0–100 scale used by MUSHRA makes it possible to express perceptible differences with a high degree of granularity, especially compared to the 0-5 modified Likert scale often used by MOS experiments. In MUSHRA, the listener is presented with the reference (labeled as such), a certain number of test samples, a hidden version of the reference, and one or more anchors (i.e. severely impaired encodings that both the experimenters and participants are supposed to immediately recognise as such; used similarly to the reference to provide a baseline demonstrating - "anchoring" - for participants the actuality of the low end of the quality scale). The recommendation specifies that a low-range and a mid-range anchor should be included in the test signals. These are typically a 7 kHz and a 3.5 kHz low-pass version of the reference. The purpose of the anchors is to calibrate the scale so that minor artifacts are not unduly penalized. This is particularly important when comparing or pooling results from different labs.

Listener behavior Both MUSHRA and ITU BS.1116 tests call for trained expert listeners who know what typical artifacts sound like and where they are likely to occur. Expert listeners also have a better internalization of the rating scale, which leads to more repeatable results than with untrained listeners. Thus, with trained listeners, fewer listeners are needed to achieve statistically significant results. It is assumed that preferences are similar for expert listeners and naive listeners, and thus, the results from expert listeners are also predictive for consumers. In agreement with this assumption Schinkel-Bielefeld et al. found no differences in the rank order between expert listeners and untrained listeners when using test signals containing only timbre and no spatial artifacts. However, Rumsey et al. showed that for signals containing spatial artifacts, expert listeners weigh spatial artifacts slightly stronger than untrained listeners, who primarily focus on timbre artifacts. In addition to this, it has been shown that expert listeners make more use of the option to listen to smaller sections of the signals under test repeatedly and perform more comparisons between the signals under test and the reference. In contrast to the naive listener who produces a preference rating, expert listeners therefore produce an audio quality rating, rating the differences between the signal under test and the uncompressed original, which is the actual goal of a MUSHRA test.

Pre- or post-screening The MUSHRA guidelines describe two major possibilities for assessing the reliability of a listener (described below). The easiest and most common is to disqualify, post-hoc, all listeners who rate the hidden reference repeat below 90 MUSHRA points for more than 15% of all test items. The hidden reference should, in the ideal case, be rated at 100 points to indicate perceptual equivalence with the original reference audio. While it can happen that the hidden reference and a high-quality signal are confused, the specification provides that a rating of lower than 90 should only be given when the listener is certain that the rated signal is different from the original reference, so a rating below 90 for the hidden reference is considered a clear and obvious listener error. The other possibility to assess a listener's performance is eGauge, a framework based on the analysis of variance (ANOVA). It computes agreement, repeatability, and discriminability, though only the latter two are recommended for pre- or post-screening. Agreement is the ANOVA of a listener's concurrence with the rest of the listeners. Repeatability examines the individual's internal reliability when rating the same test signal again in comparison to the variance of the other test signals. Discriminability analyses a sort of intertest reliability by checking that listeners can distinguish between test signals of different conditions. As eGauge requires listening to every test signal twice, its use is temporally inefficient in the immediate term relative to the prior method of post-screening listeners based on a hidden reference. eGauge does have advantages when used with a longer-term view. It negates the small chance of a complete redo in the rare case in which a sample's results lack sufficient statistical power due to an excessive failure rate discovered after the fact. Additionally, the initial inefficiency can be amortised over a series of experiments by removing the need for recruitment phases: if a listener has proven a reliable listener using eGauge, he or she can also be considered a reliable listener for future listening tests, provided the nature of the test is not substantially altered (e.g. a reliable listener for stereo tests is not necessarily equally good at perceiving artifacts in 5.1 or 22.2 configurations or potentially even mono formats).

… excerpt ends here. Continue reading the full article.

Worked examples

Example 1 — a first encounter with MUSHRA

Start with the simplest possible case. Write down what MUSHRA claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In physics, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to MUSHRA before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about MUSHRA ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of MUSHRA

In research
MUSHRA appears in physics research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses MUSHRA in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
MUSHRA is common in secondary-school and first-year university syllabi. It links to neighbouring topics ITU-R recommendations, Psychophysics, Signal processing, so understanding it makes those chapters shorter.
In everyday life
Look for MUSHRA outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study MUSHRA in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what MUSHRA means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain MUSHRA out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is MUSHRA in simple terms?

MUSHRA stands for Multiple Stimuli with Hidden Reference and Anchor and is a methodology for conducting a codec listening test to evaluate the perceived quality of the output from lossy audio compression algorithms. It is defined by ITU-R recommendation BS.1534-3.

Why does MUSHRA matter?

Because it connects several physics ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study MUSHRA?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on MUSHRA.

Tags

  • ITU-R recommendations
  • Psychophysics
  • Signal processing

Keep exploring