Crowdsourcing 2.0: Transforming QoE Testing from "Cleanup" to "Real-Time Quality Control"

Crowdsourcing 2.0: Enhancing execution speed and reliability of web-based QoE testing

2014-06-01
Bruno Gardlo, Sebastian Egger, Michael Seufert, Raimund Schatz
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces "Crowdsourcing 2.0," a framework for web-based Quality of Experience (QoE) testing that utilizes in momento (real-time) reliability verification. By replacing traditional post-hoc filtering with live monitoring and iterative task allocation, the method achieves an 86% reliability rate and reduces campaign execution time by a factor of ten.

Executive Summary

TL;DR: The transition from lab-based testing to crowdsourcing promised speed and cost-efficiency for Quality of Experience (QoE) assessment, but it has been plagued by unreliable data and "cheating" participants. This paper presents Crowdsourcing 2.0, a paradigm shift that moves from filtering bad data after a study (a posteriori) to evaluating participant reliability during the study (in momento). The result? A 10x speedup in campaign execution and a jump from 32% to 86% reliability.

Background: In the landscape of network optimization, subjective Mean Opinion Scores (MOS) are the gold standard. While traditional labs are too slow for modern agile development, this work proves that crowdsourcing can finally meet industrial reliability standards by treating participant behavior as a real-time data stream.

The Motivation: The "Cost of Distrust"

Current crowdsourcing efforts are inefficient due to a reactive workflow:

  1. Deployment: Pay for thousands of ratings.
  2. Analysis: Realize 60-70% of participants were clicking randomly or ignored the video.
  3. Filtering: Discard the majority of paid-for data.
  4. Repeat: Relaunch campaigns to fill the gap.

The authors identify Crowd Exhaustion as a hidden killer—even "reliable" workers become unreliable when faced with long, repetitive tasks (e.g., watching 20 videos). To solve this, the study asks: Can we detect a cheater within the first 60 seconds and optimize the task length based on their engagement?

Methodology: In Momento Reliability

The core innovation is a two-layered, real-time verification engine that avoids distracting questions in favor of implicit behavioral tracking.

1. The Screen Quality "Anchor"

Before the video begins, users perform a "screen quality test" disguised as calibration. They must identify contrast patterns and numbers.

  • Logic: If a user cannot identify a simple pattern or spends less than 6 seconds on the task, they receive reliability penalty points.
  • Anti-Cheat: Patterns use random movements to prevent users from sharing "correct answers" on worker forums.

2. Behavioral Telemetry

During video playback, the system monitors:

  • Browser focus (did they switch tabs?).
  • Player interactions (pausing, toggling fullscreen).
  • Total playback time.

3. The Scoring Function

Penalty points are aggregated and mapped through a Hyperbolic Tangent Function to produce a reliability percentage.

  • Dynamic Incentives: Only users who pass the initial threshold are invited to the "Extra Campaign" for more rewards, ensuring only high-quality workers proceed to more complex evaluations.

Comparison of Reliability in Studied Groups Figure: The "In Momento" approach (Study B) achieves nearly 90% reliability, far outperforming basic and even manually "Whitelisted" groups in Study A.

Experiments & Results: Speed Meets Trust

The authors conducted two large-scale studies on adaptive video streaming (85 adaptation profiles).

  • Reliability Breakthrough: Study A (Traditional) yielded only 32% usable data. Study B (In Momento) yielded 86%.
  • Efficiency Gains: Study A took 6 months to collect enough data. Study B achieved the same goal in just 25 days.
  • Cost Effective: By cutting out the waste, the cost per reliable rating dropped significantly from 0.08.

Key Result Table - Study A vs Study B Table: The quantitative performance delta. Note the massive reduction in campaign duration and the flip in reliability metrics.

Critical Insights & Conclusion

Takeaway: The "In Momento" approach solves the "Incentive-Exhaustion" paradox. By keeping tasks short (90 seconds vs. 7 minutes) and providing rapid feedback, the system respects the worker's time while safeguarding the researcher's data.

Limitations: While effective for video, this method relies heavily on Javascript-based telemetry in the browser, which can be bypassed by sophisticated bots (though much harder than standard "A-B-C" clicking).

Future Impact: This framework is a blueprint for the future of "Crowdsourcing 2.0." The move from static surveys to dynamic, behavior-aware applications will be essential as we scale QoE testing to more immersive media like 360-degree video and Cloud Gaming.

Find Similar Papers

Try Our Examples

  • Find recent studies that apply "in momento" or real-time reliability checking in large-scale human-in-the-loop (HITL) data labeling tasks beyond video QoE.
  • What are the original theoretical foundations for the "two-stage reliability framework" in crowdsourcing, and how has this paper evolved that concept into a single-stage online process?
  • Explore how behavioral tracking methods (like focus time and player interaction monitoring) have been adapted for QoE assessment in AR/VR or multi-sensory environments.
Contents
Crowdsourcing 2.0: Transforming QoE Testing from "Cleanup" to "Real-Time Quality Control"
1. Executive Summary
2. The Motivation: The "Cost of Distrust"
3. Methodology: In Momento Reliability
3.1. 1. The Screen Quality "Anchor"
3.2. 2. Behavioral Telemetry
3.3. 3. The Scoring Function
4. Experiments & Results: Speed Meets Trust
5. Critical Insights & Conclusion