CrowdMOS: Democratizing Subjective Audio Quality Assessment
CROWDMOS: An approach for crowdsourcing mean opinion score studies
The paper introduces crowdMOS, a cost-effective crowdsourcing framework for conducting Mean Opinion Score (MOS) subjective audio quality studies via Amazon Mechanical Turk. By implementing automated screening and a two-way random effects model, it achieves high-quality results comparable to laboratory studies at a fraction of the cost.
TL;DR
CrowdMOS is an open-source framework that moves Mean Opinion Score (MOS) testing from expensive labs to the internet crowd via Amazon Mechanical Turk. By using a clever Two-Way Random Effects Model to filter out bad actors and calculate honest confidence intervals, the authors prove that $11/algorithm can yield results nearly identical to professional lab studies.
The Motivation: Lab Guilt and Objective Failure
In the world of signal processing, the Mean Opinion Score (MOS) is the gold standard. However, doing it "by the book" (ITU-T P.800) is a nightmare of logistics: soundproof booths, standardized headphones, and pre-screened listeners.
Because of this, researchers often resort to Objective Measures (like PESQ). The problem? These algorithms are often "blind" to modern artifacts like jitter buffer adjustments or packet loss concealment. We need humans, but we need them to be affordable.
The Solution: CrowdMOS
The authors propose crowdMOS. It’s not just "asking people on the internet"; it’s a systematic approach to making crowdsourcing scientifically rigorous.
1. The Strategy for Human Intelligence Tasks (HITs)
To keep workers engaged and honest, the authors designed a specific UI/UX and incentive structure:
- Incentives over Barriers: Instead of a "qualification test" that scares people away, they use a bonus system based on performance and volume.
- Throughput Design: Using radio buttons and optimized layouts to ensure a worker can finish a 10-sample task in ~90 seconds.

2. The Math: Modeling the Noise
The core technical contribution is treating worker "noise" not as a nuisance, but as a statistical parameter. They use a two-way random effects model:
Where:
- : Variation in the difficulty of the sentence.
- : The specific bias/preference of the worker.
- : Pure subjective uncertainty.
By isolating these variables, they can calculate 95% Confidence Intervals (CIs) that are far more accurate than the "optimistic" CIs found in most papers that ignore worker-dependent variance.
3. The "Spam" Filter
How do you caught a worker who is just clicking "5" for everything? The system calculates a Correlation Coefficient () between a specific worker's scores and the global average. If a worker's correlation drops below 0.25, their data is nuked.
Experimental Results: The Blizzard Test
The authors put crowdMOS to the test against the Blizzard TTS Challenge. They compared 17 speech synthesizers. The results were startling:
- Repeatability: Two separate runs of crowdMOS (different days, different workers) had a 0.99 correlation.
- Accuracy: The scores matched paid UK undergraduates in a lab almost perfectly, except for two specific algorithms (T and W).

Insight: Workers with loudspeakers rated low-quality (narrowband) audio higher than lab users with headphones. This isn't a "failure" of the crowd—it actually reflects real-world usage. If your users are on laptops, the lab's "standardized" headphones might actually be giving you a false signal about what your users actually care about.
Critical Analysis & Conclusion
CrowdMOS is more than just a tool; it’s a shift in philosophy. It suggests that diversity of environment is a feature, not a bug.
Limitations:
- Expert Tests: If you are testing high-fidelity codecs where you need 20 earbuds) will fail.
- Screening Latency: You need a critical mass of scores before you can calculate the "global MOS" to filter out bad workers.
Future Work: This framework has already been extended to image quality and region-of-interest tasks, signaling a new era of "Human Computation" in signal processing.
References:
- ITU-T P.800 (The "Gold Standard")
- Amazon Mechanical Turk (The Platform)
- Open-source crowdMOS tools available via Microsoft Research.
