Crowdsourcing 2.0: Transforming QoE Testing from "Cleanup" to "Real-Time Quality Control"
Crowdsourcing 2.0: Enhancing execution speed and reliability of web-based QoE testing
This paper introduces "Crowdsourcing 2.0," a framework for web-based Quality of Experience (QoE) testing that utilizes in momento (real-time) reliability verification. By replacing traditional post-hoc filtering with live monitoring and iterative task allocation, the method achieves an 86% reliability rate and reduces campaign execution time by a factor of ten.
Executive Summary
TL;DR: The transition from lab-based testing to crowdsourcing promised speed and cost-efficiency for Quality of Experience (QoE) assessment, but it has been plagued by unreliable data and "cheating" participants. This paper presents Crowdsourcing 2.0, a paradigm shift that moves from filtering bad data after a study (a posteriori) to evaluating participant reliability during the study (in momento). The result? A 10x speedup in campaign execution and a jump from 32% to 86% reliability.
Background: In the landscape of network optimization, subjective Mean Opinion Scores (MOS) are the gold standard. While traditional labs are too slow for modern agile development, this work proves that crowdsourcing can finally meet industrial reliability standards by treating participant behavior as a real-time data stream.
The Motivation: The "Cost of Distrust"
Current crowdsourcing efforts are inefficient due to a reactive workflow:
- Deployment: Pay for thousands of ratings.
- Analysis: Realize 60-70% of participants were clicking randomly or ignored the video.
- Filtering: Discard the majority of paid-for data.
- Repeat: Relaunch campaigns to fill the gap.
The authors identify Crowd Exhaustion as a hidden killer—even "reliable" workers become unreliable when faced with long, repetitive tasks (e.g., watching 20 videos). To solve this, the study asks: Can we detect a cheater within the first 60 seconds and optimize the task length based on their engagement?
Methodology: In Momento Reliability
The core innovation is a two-layered, real-time verification engine that avoids distracting questions in favor of implicit behavioral tracking.
1. The Screen Quality "Anchor"
Before the video begins, users perform a "screen quality test" disguised as calibration. They must identify contrast patterns and numbers.
- Logic: If a user cannot identify a simple pattern or spends less than 6 seconds on the task, they receive reliability penalty points.
- Anti-Cheat: Patterns use random movements to prevent users from sharing "correct answers" on worker forums.
2. Behavioral Telemetry
During video playback, the system monitors:
- Browser focus (did they switch tabs?).
- Player interactions (pausing, toggling fullscreen).
- Total playback time.
3. The Scoring Function
Penalty points are aggregated and mapped through a Hyperbolic Tangent Function to produce a reliability percentage.
- Dynamic Incentives: Only users who pass the initial threshold are invited to the "Extra Campaign" for more rewards, ensuring only high-quality workers proceed to more complex evaluations.
Figure: The "In Momento" approach (Study B) achieves nearly 90% reliability, far outperforming basic and even manually "Whitelisted" groups in Study A.
Experiments & Results: Speed Meets Trust
The authors conducted two large-scale studies on adaptive video streaming (85 adaptation profiles).
- Reliability Breakthrough: Study A (Traditional) yielded only 32% usable data. Study B (In Momento) yielded 86%.
- Efficiency Gains: Study A took 6 months to collect enough data. Study B achieved the same goal in just 25 days.
- Cost Effective: By cutting out the waste, the cost per reliable rating dropped significantly from 0.08.
Table: The quantitative performance delta. Note the massive reduction in campaign duration and the flip in reliability metrics.
Critical Insights & Conclusion
Takeaway: The "In Momento" approach solves the "Incentive-Exhaustion" paradox. By keeping tasks short (90 seconds vs. 7 minutes) and providing rapid feedback, the system respects the worker's time while safeguarding the researcher's data.
Limitations: While effective for video, this method relies heavily on Javascript-based telemetry in the browser, which can be bypassed by sophisticated bots (though much harder than standard "A-B-C" clicking).
Future Impact: This framework is a blueprint for the future of "Crowdsourcing 2.0." The move from static surveys to dynamic, behavior-aware applications will be essential as we scale QoE testing to more immersive media like 360-degree video and Cloud Gaming.
