Crowdsourcing Video Quality: Is the Lab Environment Obsolete for Internet Video?
Assessing internet video quality using crowdsourcing
The paper introduces a scalable, HTML5-based subjective video quality evaluation system integrated with crowdsourcing platforms like Amazon Mechanical Turk and CrowdFlower. By utilizing AVC/H.264 bitstreams with known Mean Opinion Scores (MOS), the study validates that crowdsourced evaluations can effectively replicate formal MPEG laboratory results for Internet-targeted video bitrates.
TL;DR
This research investigates whether the "crowd" can replace expensive laboratory testing for video quality assessment. By building an HTML5-based evaluation system and deploying it on Amazon Mechanical Turk and CrowdFlower, the researchers proved that crowdsourced ratings are consistent, reproducible, and correlate strongly with formal MPEG standards, albeit with a consistent "leniency" offset due to real-world viewing conditions.
The "Golden Standard" Problem
In the world of video compression, the Mean Opinion Score (MOS) remains the ultimate benchmark. While objective metrics like PSNR or SSIM are easy to calculate, they often fail to capture the nuances of human perception. Historically, getting a MOS required the ITU-R Bt.500 ritual: a dark room, specialized monitors, calibrated distances, and a group of "vetted" human subjects.
The authors argue this method is fundamentally flawed for the modern age:
- Scalability: You can't run a lab test every time you tweak a minor encoder setting.
- Realism: People don't watch YouTube in a sterile lab; they watch it on laptops and phones in coffee shops or bedrooms.
Methodology: Bringing the Lab to the Browser
The researchers developed a system that uses HTML5 and JavaScript to present pairs of videos (Reference vs. Compressed) to workers across the globe.
The Workflow Architecture
Participants take a survey on habits, watch 10-second clips, and rate them on a 0-10 scale. To ensure data integrity, the system includes:
- Resolution Querying: Ensuring the video is never upscaled or downscaled by the browser.
- Gold Questions: Random content-based questions (e.g., "What color was the car?") to identify and ban automated bots or inattentive workers.
- R5 Reference Strategy: Since streaming "lossless" video to a crowd is impossible due to bandwidth, they used high-bitrate "beta-anchors" as the 10/10 reference.
Figure 1: The assessment workflow from instructions to voting.
Key Findings: The "Leniency" Offset
The most striking result was the comparison between the Paid Crowd, Volunteer University Crowd, and the MPEG Lab Anchors.
- Consistency: The Paid Crowd and Volunteers produced nearly identical results. Crowdsourcing is reproducible.
- The 3.2 Point Gap: Crowdsourced scores were consistently higher than Lab scores by an average of 3.2 points.
Why the higher scores?
The authors suggest that because the crowd watches video on standard monitors in non-ideal lighting, they are more tolerant of compression artifacts. In professional labs, subjects are trained to hunt for "mosquito noise" and "blocking." In the real world, if the video looks "good enough," users rate it highly.
Figure 2: The bitstreams used for testing (Classes C, D, and E).
Deep Insight: Efficiency vs. Accuracy
The experiment highlights a crucial industry insight: Video providers might be over-encoding. If the crowd perceives a lower bitrate stream as "High Quality" while a lab expert calls it "Poor," the business logic dictates following the crowd. This saves bandwidth costs without sacrificing perceived user satisfaction.
Figure 3: Qualitative results across different sequences (BasketBallDrill, BQmall, etc.).
Critical Analysis & Conclusion
While the study successfully validates crowdsourcing as a tool, it also identifies a "systematic shift" in ratings that needs further modeling. The 3.2-point difference suggests we cannot use Crowdsourced MOS and Lab MOS interchangeably without a correction factor.
Takeaway: Crowdsourcing is no longer just a "cheap trick" for data labeling; it is a scientifically valid method for assessing video quality that more accurately reflects the "Wild West" of the Internet than any sterile laboratory could.
Future Work: The authors aim to investigate how specific habits—like being an "action gamer"—influence a person's sensitivity to frame-rate drops and compression artifacts.
