Lab Experiment vs. Crowdsourcing: Scaling the Measurement of User Experience

Lab experiment vs. crowdsourcing: a comparative user study on Skype call quality

2013-11-13
Yu-Chuan Yen, Cing-Yu Chu, Su-Ling Yeh, Hao-Hua Chu, Polly Huang, Polly Huang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comparative study between traditional laboratory experiments and crowdsourcing (via Amazon Mechanical Turk) for evaluating the Quality of Experience (QoE) of Skype calls. Using the SILK codec across various bitrates, the authors demonstrate that crowdsourcing provides comparable Mean Opinion Score (MOS) results to lab settings while significantly outperforming them in efficiency and participant diversity.

TL;DR

Is a controlled lab environment truly necessary to measure how users perceive call quality? This study says "No." By comparing Skype call quality evaluations conducted in a university lab versus Amazon Mechanical Turk (MTurk), researchers found that crowdsourcing is not only faster (1 hour vs. weeks) and cheaper (10% of the cost) but also provides a more diverse and statistically robust dataset without sacrificing accuracy.

Background: The Scalability Crisis in QoE

In the world of network engineering, there is a fundamental split between Quality of Service (QoS)—the objective metrics like packet loss and jitter—and Quality of Experience (QoE)—how the user actually feels.

While we can measure QoS with scripts, QoE has traditionally required human subjects sitting in a quiet room with expensive headphones. This "Lab Paradigm" creates a bottleneck: it’s too slow to keep up with agile protocol development and too demographic-limited (mostly young male engineering students) to represent the global internet population.

Methodology: Bridging the Gap with MTurk

The researchers set out to validate whether the "wild west" of crowdsourcing could match the "sterile" lab. They tested 9 different bitrates of the Skype SILK codec (ranging from 5.6 kbps to 40.6 kbps).

The "Cheat-Proof" Mechanism

The biggest criticism of crowdsourcing is data fidelity—how do you know the worker isn't just clicking random buttons for a dollar? The authors used a clever solution:

  1. Hidden Reference Tracks: They randomly inserted high-quality audio samples.
  2. Rejection Criteria: If a user gave a high-quality reference a low score (MOS ≤ 3), their entire submission was discarded.

MOS-bitrate Relationship Figure 1: The theoretical logarithmic relationship between Bitrate and Mean Opinion Score (MOS).

Experiments & Results: Efficiency vs. Validity

1. The Cost-Time Revolution

The table below highlights the staggering efficiency of crowdsourcing. While the lab experiment took weeks of scheduling and 5 hours of active supervision for just 30 participants, the MTurk study collected 60 valid samples in just 1 hour.

MetricLaboratory ExperimentAmazon MTurk
Supervising Time8 mins/submission1 min/submission
Data Collection Time~Weeks (Active 5 hrs)1 Hour
Total Cost~$670 (iPhone Lottery)1/user + fee)

2. Statistical Convergence

The core finding of the paper is the convergence of results. As shown in the comparison graph, the MOS curves for both groups follow the same logarithmic trend. Even though MTurk users used varied devices (laptop speakers, in-ear headphones, etc.), the average perception of quality remained consistent with the controlled lab environment.

Lab and MTurk Result Comparison Figure 2: Comparing MOS results highlights that MTurk (red) closely tracks the Lab experiment (blue).

Deep Dive: Does Device or Age Matter?

Because MTurk allowed for a larger and more diverse pool, the authors performed "fine-grained" analysis that the lab study couldn't support:

  • Gender & Age: Neither gender nor age significantly altered the perception of call quality. The logarithmic sensitivity to bitrate is a near-universal human trait.
  • Device Factor: Interestingly, while people using headphones vs. speakers gave similar average scores, headphone users produced a "smoother" curve. Methodologically, this suggests that headphones make both the beauty and the flaws of a codec clearer, reducing noise in the data.

Critical Insight & Conclusion

This paper serves as a green light for the networking community to embrace crowdsourcing. The "Gold Standard" of the laboratory is often an unnecessary bottleneck for audio codec evaluation.

Key Takeaways for Future Researchers:

  • Diversity is a Feature, not a Bug: The varied environments of MTurk users actually make the QoE results more "real-world" than a silent lab.
  • Validation is Key: Crowdsourcing is only as good as your "cheat-proof" tests. Planting known-quality samples is non-negotiable for maintaining High-Quality Data.

By moving QoE out of the lab, we can finally scale our understanding of user experience to the scale of the Internet itself.

Find Similar Papers

Try Our Examples

  • Find recent papers that compare crowdsourcing and laboratory environments for Video QoE or Virtual Reality (VR) subjective assessments.
  • Which study first proposed the Weber-Fechner Law as the theoretical basis for the logarithmic relationship between QoS and QoE in telecommunications?
  • Explore how advanced "cheat-proof" or "gold standard" mechanisms in crowdsourcing have evolved beyond simple hidden reference tracks to filter out sophisticated bot responses.
Contents
Lab Experiment vs. Crowdsourcing: Scaling the Measurement of User Experience
1. TL;DR
2. Background: The Scalability Crisis in QoE
3. Methodology: Bridging the Gap with MTurk
3.1. The "Cheat-Proof" Mechanism
4. Experiments & Results: Efficiency vs. Validity
4.1. 1. The Cost-Time Revolution
4.2. 2. Statistical Convergence
5. Deep Dive: Does Device or Age Matter?
6. Critical Insight & Conclusion