Crowdsourcing Video Quality: Is the Lab Environment Obsolete for Internet Video?

Assessing internet video quality using crowdsourcing

2013-10-17
Oscar Figuerola Salas, Velibor Adzic, Akash Shah, Hari Kalva
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a scalable, HTML5-based subjective video quality evaluation system integrated with crowdsourcing platforms like Amazon Mechanical Turk and CrowdFlower. By utilizing AVC/H.264 bitstreams with known Mean Opinion Scores (MOS), the study validates that crowdsourced evaluations can effectively replicate formal MPEG laboratory results for Internet-targeted video bitrates.

TL;DR

This research investigates whether the "crowd" can replace expensive laboratory testing for video quality assessment. By building an HTML5-based evaluation system and deploying it on Amazon Mechanical Turk and CrowdFlower, the researchers proved that crowdsourced ratings are consistent, reproducible, and correlate strongly with formal MPEG standards, albeit with a consistent "leniency" offset due to real-world viewing conditions.

The "Golden Standard" Problem

In the world of video compression, the Mean Opinion Score (MOS) remains the ultimate benchmark. While objective metrics like PSNR or SSIM are easy to calculate, they often fail to capture the nuances of human perception. Historically, getting a MOS required the ITU-R Bt.500 ritual: a dark room, specialized monitors, calibrated distances, and a group of "vetted" human subjects.

The authors argue this method is fundamentally flawed for the modern age:

  1. Scalability: You can't run a lab test every time you tweak a minor encoder setting.
  2. Realism: People don't watch YouTube in a sterile lab; they watch it on laptops and phones in coffee shops or bedrooms.

Methodology: Bringing the Lab to the Browser

The researchers developed a system that uses HTML5 and JavaScript to present pairs of videos (Reference vs. Compressed) to workers across the globe.

The Workflow Architecture

Participants take a survey on habits, watch 10-second clips, and rate them on a 0-10 scale. To ensure data integrity, the system includes:

  • Resolution Querying: Ensuring the video is never upscaled or downscaled by the browser.
  • Gold Questions: Random content-based questions (e.g., "What color was the car?") to identify and ban automated bots or inattentive workers.
  • R5 Reference Strategy: Since streaming "lossless" video to a crowd is impossible due to bandwidth, they used high-bitrate "beta-anchors" as the 10/10 reference.

System Structure Figure 1: The assessment workflow from instructions to voting.

Key Findings: The "Leniency" Offset

The most striking result was the comparison between the Paid Crowd, Volunteer University Crowd, and the MPEG Lab Anchors.

  1. Consistency: The Paid Crowd and Volunteers produced nearly identical results. Crowdsourcing is reproducible.
  2. The 3.2 Point Gap: Crowdsourced scores were consistently higher than Lab scores by an average of 3.2 points.

Why the higher scores?

The authors suggest that because the crowd watches video on standard monitors in non-ideal lighting, they are more tolerant of compression artifacts. In professional labs, subjects are trained to hunt for "mosquito noise" and "blocking." In the real world, if the video looks "good enough," users rate it highly.

Table of Video Classes Figure 2: The bitstreams used for testing (Classes C, D, and E).

Deep Insight: Efficiency vs. Accuracy

The experiment highlights a crucial industry insight: Video providers might be over-encoding. If the crowd perceives a lower bitrate stream as "High Quality" while a lab expert calls it "Poor," the business logic dictates following the crowd. This saves bandwidth costs without sacrificing perceived user satisfaction.

MOS Results Comparison Figure 3: Qualitative results across different sequences (BasketBallDrill, BQmall, etc.).

Critical Analysis & Conclusion

While the study successfully validates crowdsourcing as a tool, it also identifies a "systematic shift" in ratings that needs further modeling. The 3.2-point difference suggests we cannot use Crowdsourced MOS and Lab MOS interchangeably without a correction factor.

Takeaway: Crowdsourcing is no longer just a "cheap trick" for data labeling; it is a scientifically valid method for assessing video quality that more accurately reflects the "Wild West" of the Internet than any sterile laboratory could.

Future Work: The authors aim to investigate how specific habits—like being an "action gamer"—influence a person's sensitivity to frame-rate drops and compression artifacts.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize crowdsourcing for Quality of Experience (QoE) assessment in 4K or VR/AR video streaming.
  • Which study first introduced the "CrowdMOS" metric, and how does this paper's HTML5 system implementation differ from that methodology?
  • How have modern deep learning-based objective metrics, such as Netflix’s VMAF, been validated against crowdsourced subjective datasets like the ones discussed in this paper?
Contents
Crowdsourcing Video Quality: Is the Lab Environment Obsolete for Internet Video?
1. TL;DR
2. The "Golden Standard" Problem
3. Methodology: Bringing the Lab to the Browser
3.1. The Workflow Architecture
4. Key Findings: The "Leniency" Offset
4.1. Why the higher scores?
5. Deep Insight: Efficiency vs. Accuracy
6. Critical Analysis & Conclusion