Beyond the Average: A Probabilistic Shift in Subjective Quality Assessment

A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from Crowdsourcing

2020-10-12
Jing Li, Suiyi Ling, Junle Wang, Patrick Le Callet
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a Generic Probabilistic Model (GPM) for inferring ground-truth visual quality from noisy crowdsourced data. By modeling perceptual quality as an ordinal categorical distribution and utilizing an EM algorithm to distinguish serious annotators from spammers, it achieves state-of-the-art performance in Mean Opinion Score (MOS) recovery.

TL;DR

Researchers from Alibaba, Tencent, and Nantes University have developed a Generic Probabilistic Model (GPM) to clean noisy crowdsourced data. By shifting the definition of "ground truth" from a single Mean Opinion Score (MOS) to a categorical distribution, the model better reflects human subjectivity and effectively identifies spammers, significantly outperforming traditional ITU standards and existing SOTA truth-inference algorithms.

The Problem: The "Average" User Doesn't Exist

In the world of Video Quality Assessment (VQA), the gold standard is the Mean Opinion Score (MOS). Traditionally, we gather people in a lab, ask them to rate a video from 1 to 5, and average the results.

However, crowdsourcing platforms like Amazon Mechanical Turk introduce chaos. Annotators might be "picky," "optimistic," or "spammers" who click randomly for a quick payout. Existing methods often:

  1. Discard too much data: Standard ITU-R BT.500 rejection rules are often too aggressive.
  2. Assume Gaussianity: They assume human errors follow a "Bell Curve," which research shows is often false for perceptual tasks.

Methodology: Modeling the Soul of the Annotator

The brilliance of the GPM lies in its Probabilistic Graphical Model structure. It doesn't ask "What is the correct score?" but rather "What is the probability that a reliable observer chooses score ?"

1. The Ground Truth Distribution

Instead of a single value, the ground truth for an object is a vector , representing an ordinal categorical distribution. This respects the fact that two reliable humans might legitimately disagree on whether a video is "Good" or "Fair."

2. The Latent Switch (Reliability vs. Irregularity)

For every rating, the model assumes a latent variable (reliability).

  • If : The annotator is acting seriously, guided by the ground truth .
  • If : The annotator is acting irregularly, guided by a spamming profile .

Overall Architecture Figure 1: The Graphical Model showing the interaction between ground truth (), reliability (), and irregular behavior ().

The authors solve this using the Expectation-Maximization (EM) algorithm, iteratively refining the estimate of who is a spammer and what the true quality distribution is.

Experimental Battleground

The model was put to the test against industry heavyweights like Dawid-Skene (D&S) and REML.

Robustness to Spammers

In simulations using real HDTV and UHD video datasets, the authors increased the "spamminess ratio" from 20% to 80%. While traditional MOS and BT.500 rejection crashed as noise increased, GPM remained remarkably stable.

Performance Comparison Figure 2: Impact of spammer percentage. GPM (orange line) maintains the highest ROCC and lowest RMSE compared to all baselines.

Visualizing Behavior

One of the most practical outputs of this research is the ability to "fingerprint" annotator behavior. The model can identify:

  • Optimistic Bias: Always rating higher than the consensus.
  • Picky Bias: Always rating lower.
  • Random Spammers: No correlation with actual quality.

Annotator Behaviors Figure 5: Different archetypes of annotators identified by the GPM model.

Deep Insights & Future Outlook

The GPM model proves that subjectivity is not noise—it's data. By modeling the distribution of opinions, we gain a richer understanding of media quality than a simple average ever could.

Limitations: While the categorical distribution is powerful, the model currently uses a "loose" uniform distribution for irregular behavior. Future work could refine this to specifically target "malicious adversaries" who intentionally invert scores to sabotage datasets.

Takeaway for Practitioners: If you are building a dataset via crowdsourcing, stop using simple Mean Voting. Adopting a probabilistic approach like GPM allows you to keep more data while gaining higher precision, even with a high percentage of unreliable workers.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Bayesian Crowdsourcing or Truth Discovery methods specifically for multimedia Quality of Experience (QoE) assessment.
  • What is the origin of the Dawid-Skene model, and how have modern deep learning approaches integrated its expectation-maximization logic for noisy label handling?
  • Explore if the proposed ordinal categorical distribution approach has been applied to subjective datasets in other domains like Medical Imaging or Affective Computing.
Contents
Beyond the Average: A Probabilistic Shift in Subjective Quality Assessment
1. TL;DR
2. The Problem: The "Average" User Doesn't Exist
3. Methodology: Modeling the Soul of the Annotator
3.1. 1. The Ground Truth Distribution
3.2. 2. The Latent Switch (Reliability vs. Irregularity)
4. Experimental Battleground
4.1. Robustness to Spammers
4.2. Visualizing Behavior
5. Deep Insights & Future Outlook