Lab Experiment vs. Crowdsourcing: Scaling the Measurement of User Experience
Lab experiment vs. crowdsourcing: a comparative user study on Skype call quality
This paper presents a comparative study between traditional laboratory experiments and crowdsourcing (via Amazon Mechanical Turk) for evaluating the Quality of Experience (QoE) of Skype calls. Using the SILK codec across various bitrates, the authors demonstrate that crowdsourcing provides comparable Mean Opinion Score (MOS) results to lab settings while significantly outperforming them in efficiency and participant diversity.
TL;DR
Is a controlled lab environment truly necessary to measure how users perceive call quality? This study says "No." By comparing Skype call quality evaluations conducted in a university lab versus Amazon Mechanical Turk (MTurk), researchers found that crowdsourcing is not only faster (1 hour vs. weeks) and cheaper (10% of the cost) but also provides a more diverse and statistically robust dataset without sacrificing accuracy.
Background: The Scalability Crisis in QoE
In the world of network engineering, there is a fundamental split between Quality of Service (QoS)—the objective metrics like packet loss and jitter—and Quality of Experience (QoE)—how the user actually feels.
While we can measure QoS with scripts, QoE has traditionally required human subjects sitting in a quiet room with expensive headphones. This "Lab Paradigm" creates a bottleneck: it’s too slow to keep up with agile protocol development and too demographic-limited (mostly young male engineering students) to represent the global internet population.
Methodology: Bridging the Gap with MTurk
The researchers set out to validate whether the "wild west" of crowdsourcing could match the "sterile" lab. They tested 9 different bitrates of the Skype SILK codec (ranging from 5.6 kbps to 40.6 kbps).
The "Cheat-Proof" Mechanism
The biggest criticism of crowdsourcing is data fidelity—how do you know the worker isn't just clicking random buttons for a dollar? The authors used a clever solution:
- Hidden Reference Tracks: They randomly inserted high-quality audio samples.
- Rejection Criteria: If a user gave a high-quality reference a low score (MOS ≤ 3), their entire submission was discarded.
Figure 1: The theoretical logarithmic relationship between Bitrate and Mean Opinion Score (MOS).
Experiments & Results: Efficiency vs. Validity
1. The Cost-Time Revolution
The table below highlights the staggering efficiency of crowdsourcing. While the lab experiment took weeks of scheduling and 5 hours of active supervision for just 30 participants, the MTurk study collected 60 valid samples in just 1 hour.
| Metric | Laboratory Experiment | Amazon MTurk |
|---|---|---|
| Supervising Time | 8 mins/submission | 1 min/submission |
| Data Collection Time | ~Weeks (Active 5 hrs) | 1 Hour |
| Total Cost | ~$670 (iPhone Lottery) | 1/user + fee) |
2. Statistical Convergence
The core finding of the paper is the convergence of results. As shown in the comparison graph, the MOS curves for both groups follow the same logarithmic trend. Even though MTurk users used varied devices (laptop speakers, in-ear headphones, etc.), the average perception of quality remained consistent with the controlled lab environment.
Figure 2: Comparing MOS results highlights that MTurk (red) closely tracks the Lab experiment (blue).
Deep Dive: Does Device or Age Matter?
Because MTurk allowed for a larger and more diverse pool, the authors performed "fine-grained" analysis that the lab study couldn't support:
- Gender & Age: Neither gender nor age significantly altered the perception of call quality. The logarithmic sensitivity to bitrate is a near-universal human trait.
- Device Factor: Interestingly, while people using headphones vs. speakers gave similar average scores, headphone users produced a "smoother" curve. Methodologically, this suggests that headphones make both the beauty and the flaws of a codec clearer, reducing noise in the data.
Critical Insight & Conclusion
This paper serves as a green light for the networking community to embrace crowdsourcing. The "Gold Standard" of the laboratory is often an unnecessary bottleneck for audio codec evaluation.
Key Takeaways for Future Researchers:
- Diversity is a Feature, not a Bug: The varied environments of MTurk users actually make the QoE results more "real-world" than a silent lab.
- Validation is Key: Crowdsourcing is only as good as your "cheat-proof" tests. Planting known-quality samples is non-negotiable for maintaining High-Quality Data.
By moving QoE out of the lab, we can finally scale our understanding of user experience to the scale of the Internet itself.
