Beyond 0 and 1: Crowdsourcing Relevance Magnitudes for Smarter IR Evaluation
On crowdsourcing relevance magnitudes for information retrieval evaluation
This paper introduces "Magnitude Estimation" (ME), a psychophysical scaling technique, to the field of Information Retrieval (IR) for judging document relevance. Unlike traditional ordinal scales (e.g., 0-3), ME allows users to assign continuous numbers representing perceived relevance ratios, achieving a large-scale evaluation using crowdsourced data across 18 TREC topics and 50,000+ judgments.
TL;DR
Is a "highly relevant" document exactly twice as good as a "marginally relevant" one? Most IR metrics assume so, but this paper challenges that dogma. By applying Magnitude Estimation (ME)—a psychophysical technique from the 1950s—to 50,000+ crowdsourced judgments, the authors prove that relevance is a continuous ratio, not a rigid category. Their findings show that current evaluation metrics (nDCG, ERR) might be misrepresenting how users actually value information.
The Problem: The Tyranny of the Ordinal Scale
For decades, Information Retrieval (IR) has lived in a world of boxes. We label documents as Irrelevant (0), Relevant (1), or Highly Relevant (2). These ordinal scales are convenient but mathematically flawed for many operations.
The core issues are:
- Ambiguous Boundaries: Where does "Marginal" end and "Relevant" begin?
- Undefined Distances: Is the jump from level 1 to 2 the same as 0 to 1?
- Loss of Detail: We lose the nuances of how much better one document is than another.
Methodology: Bringing Psychophysics to the Crowd
The authors turned to Magnitude Estimation. Instead of picking a category, a user is shown a document and asked to assign any positive number to its relevance. If the next document is twice as relevant, they give it a number twice as large.
To ensure high-quality data from crowdsourcing (CrowdFlower), they utilized a clever 4-step pipeline:
- The Line Test: Users first estimated the length of lines to prove they understood ratio scaling.
- Topic Anchors: Every user saw a known "High" (Hk) and "Low" (Nk) document. If they rated Nk higher than Hk, their data was tossed.
- Geometric Normalization: Since one user might use a scale of 1-10 and another 100-1000, the authors used geometric averaging to normalize the data while keeping the ratios intact.
The raw scores follow a log-normal distribution, typical of human sensory perception.
Key Insights: System Rankings in Turmoil
The most shocking result came when the authors recalculated the rankings of actual TREC-8 search systems.
When using ME scores as the "gain" (the value a user gets from a document), the system rankings shifted dramatically.
- Kendall’s correlation between Magnitude-based nDCG and Category-based nDCG was only 0.677.
- For the ERR metric, the correlation was even lower, suggesting that the top-performing systems identified by traditional metrics might not be the best from a user-perception standpoint.
Discrepancy in system effectiveness: The red dots show systems that are statistically "the best." Note the lack of overlap when using different relevance scales.
Why Does This Happen?
The study found that users fall into two camps:
- Narrow Scalers: Users who perceive small differences between "good" and "great" documents.
- Wide Scalers: Users who perceive "Highly Relevant" documents as being thousands of times more valuable than marginally relevant ones.
Traditional metrics like nDCG usually use a linear (0, 1, 2, 3) or exponential () gain profile. This paper proves that neither profile fits all users. The "average" user is closer to linear, but a significant portion of the population is completely ignored by these fixed assumptions.
Critical Analysis & Conclusion
Takeaway
This paper is a wakeup call for IR researchers. It demonstrates that Magnitude Estimation is not only viable for crowdsourcing but captures a dimension of "user satisfaction" that current benchmarks ignore. It suggests that our "SOTA" rankings are built on the shaky assumption of a uniform user perception.
Limitations
- Complexity: ME is harder for users than simple clicking bits; it requires 78 seconds per document vs. typical rapid-fire labeling.
- No Zero: ME forbids zero/negative numbers, which makes "truly irrelevant" documents tricky to handle (users often use tiny decimals).
Future Work
The next step is applying ME to novelty and diversity. Perhaps a document isn't just "relevant," but its magnitude of relevance drops once you've seen a similar one. Magnitude Estimation provides the mathematical ratio scale needed to solve these complex interactions.
Senior Editor's Note: This work bridges the gap between psychophysics and data science, reminding us that behind every "label" is a human sensation that a single integer can rarely capture.
