Beyond 0 and 1: Crowdsourcing Relevance Magnitudes for Smarter IR Evaluation

On crowdsourcing relevance magnitudes for information retrieval evaluation

2024-11-02
Eddy Maddalena (19998288), Stefano Mizzaro (8223846), Falk Scholer (18029842), Andrew Turpin (4287529)
Summary
Problem
Method
Results
Takeaways

This paper introduces "Magnitude Estimation" (ME), a psychophysical scaling technique, to the field of Information Retrieval (IR) for judging document relevance. Unlike traditional ordinal scales (e.g., 0-3), ME allows users to assign continuous numbers representing perceived relevance ratios, achieving a large-scale evaluation using crowdsourced data across 18 TREC topics and 50,000+ judgments.

TL;DR

Is a "highly relevant" document exactly twice as good as a "marginally relevant" one? Most IR metrics assume so, but this paper challenges that dogma. By applying Magnitude Estimation (ME)—a psychophysical technique from the 1950s—to 50,000+ crowdsourced judgments, the authors prove that relevance is a continuous ratio, not a rigid category. Their findings show that current evaluation metrics (nDCG, ERR) might be misrepresenting how users actually value information.

The Problem: The Tyranny of the Ordinal Scale

For decades, Information Retrieval (IR) has lived in a world of boxes. We label documents as Irrelevant (0), Relevant (1), or Highly Relevant (2). These ordinal scales are convenient but mathematically flawed for many operations.

The core issues are:

  1. Ambiguous Boundaries: Where does "Marginal" end and "Relevant" begin?
  2. Undefined Distances: Is the jump from level 1 to 2 the same as 0 to 1?
  3. Loss of Detail: We lose the nuances of how much better one document is than another.

Methodology: Bringing Psychophysics to the Crowd

The authors turned to Magnitude Estimation. Instead of picking a category, a user is shown a document and asked to assign any positive number to its relevance. If the next document is twice as relevant, they give it a number twice as large.

To ensure high-quality data from crowdsourcing (CrowdFlower), they utilized a clever 4-step pipeline:

  • The Line Test: Users first estimated the length of lines to prove they understood ratio scaling.
  • Topic Anchors: Every user saw a known "High" (Hk) and "Low" (Nk) document. If they rated Nk higher than Hk, their data was tossed.
  • Geometric Normalization: Since one user might use a scale of 1-10 and another 100-1000, the authors used geometric averaging to normalize the data while keeping the ratios intact.

Experimental Distribution of Scores The raw scores follow a log-normal distribution, typical of human sensory perception.

Key Insights: System Rankings in Turmoil

The most shocking result came when the authors recalculated the rankings of actual TREC-8 search systems.

When using ME scores as the "gain" (the value a user gets from a document), the system rankings shifted dramatically.

  • Kendall’s correlation between Magnitude-based nDCG and Category-based nDCG was only 0.677.
  • For the ERR metric, the correlation was even lower, suggesting that the top-performing systems identified by traditional metrics might not be the best from a user-perception standpoint.

ME vs TREC System Rankings Discrepancy in system effectiveness: The red dots show systems that are statistically "the best." Note the lack of overlap when using different relevance scales.

Why Does This Happen?

The study found that users fall into two camps:

  1. Narrow Scalers: Users who perceive small differences between "good" and "great" documents.
  2. Wide Scalers: Users who perceive "Highly Relevant" documents as being thousands of times more valuable than marginally relevant ones.

Traditional metrics like nDCG usually use a linear (0, 1, 2, 3) or exponential () gain profile. This paper proves that neither profile fits all users. The "average" user is closer to linear, but a significant portion of the population is completely ignored by these fixed assumptions.

Critical Analysis & Conclusion

Takeaway

This paper is a wakeup call for IR researchers. It demonstrates that Magnitude Estimation is not only viable for crowdsourcing but captures a dimension of "user satisfaction" that current benchmarks ignore. It suggests that our "SOTA" rankings are built on the shaky assumption of a uniform user perception.

Limitations

  • Complexity: ME is harder for users than simple clicking bits; it requires 78 seconds per document vs. typical rapid-fire labeling.
  • No Zero: ME forbids zero/negative numbers, which makes "truly irrelevant" documents tricky to handle (users often use tiny decimals).

Future Work

The next step is applying ME to novelty and diversity. Perhaps a document isn't just "relevant," but its magnitude of relevance drops once you've seen a similar one. Magnitude Estimation provides the mathematical ratio scale needed to solve these complex interactions.


Senior Editor's Note: This work bridges the gap between psychophysics and data science, reminding us that behind every "label" is a human sensation that a single integer can rarely capture.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize psychophysical scaling or magnitude estimation for evaluating Large Language Model (LLM) responses or search result quality.
  • Search for the original psychophysics papers by Stanley Stevens that established Magnitude Estimation as a ratio-scaling technique and compare its application in HCI.
  • Which modern Information Retrieval metrics have attempted to incorporate user subjectivity or individual gain profiles beyond the static weights of nDCG and ERR?
Contents
Beyond 0 and 1: Crowdsourcing Relevance Magnitudes for Smarter IR Evaluation
1. TL;DR
2. The Problem: The Tyranny of the Ordinal Scale
3. Methodology: Bringing Psychophysics to the Crowd
4. Key Insights: System Rankings in Turmoil
5. Why Does This Happen?
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Work