Validation of Smets’ Hypothesis: Why Imprecise Humans are More Certain
Validation of Smets’ Hypothesis in the Crowdsourcing Environment
This paper experimentally validates Smets' Hypothesis using two crowdsourcing campaigns for bird image annotation. It demonstrates that allowing users to provide imprecise answers (selecting multiple options) significantly increases their self-reported certainty, achieving a 90% accuracy rate using belief function aggregation.
TL;DR
Is a person more certain when they are forced to pick one answer, or when they can pick several? This paper validates a 30-year-old hypothesis by Philippe Smets: imprecision breeds certainty. By allowing crowdsourcing workers to select multiple bird species in an annotation task, the researchers found that workers became significantly more confident, ultimately leading to a 6% boost in final accuracy (reaching 90%) when aggregated using the Theory of Belief Functions.
Problem & Motivation: The Tyranny of the Forced Choice
In standard crowdsourcing (like Amazon Mechanical Turk), we treat humans like binary switches. We ask: "Is this a Robin or a Sparrow?" Even if the worker is 51/49% unsure, they must flip a coin and pick one.
This "forced precision" creates a paradox. The worker is precise (one answer) but highly uncertain. Philippe Smets hypothesized in 1997 that:
- H1: The more imprecise a person is, the more certain they are.
- H2: The more precise a person is, the less certain they are.
Until now, this remained a theoretical intuition. This paper seeks the first experimental validation in a real-world data collection environment.
Methodology: Testing the "Imprecise" Edge
The researchers deployed two distinct campaigns on the Crowdpanel platform, each involving 100 users and 50 bird photos:
- Experiment 1 (Precise): Users must select exactly one bird name (Radio buttons).
- Experiment 2 (Imprecise): Users can select between 1 and 5 names (Checkboxes).
In both cases, users rated their certainty on a scale of 0 to 6.
Modeling via Belief Functions
The core mathematical engine is the Theory of Belief Functions. Instead of simple probabilities, the authors use mass functions ():
- Imprecision is captured by the set (e.g., {Robin, Sparrow}).
- Uncertainty is captured by the weight assigned to that set, derived from the user's certainty score.
- Ignorance is the remaining mass assigned to the whole frame .
The Simple Mass Function used to blend uncertainty and imprecision.
Experiments & Results: The "Certainty Gap"
The data confirms the hypothesis with striking clarity. As shown in the comparative analysis, users who were allowed to be imprecise (Experiment 2) reported consistently higher certainty scores across all levels of task difficulty.
Key Findings:
- Higher Confidence: Average certainty in the imprecise group was 29.68% higher than the precise group.
- Accuracy Gains: When these imprecise, high-certainty answers were aggregated using Dempster's Rule of Combination, the final accuracy reached 90%.
- Superior Aggregation: Even if using simple Majority Voting, the imprecise group reached 83% accuracy, far outperforming the 70% achieved by the precise group in similar voting conditions.
Figure (b) shows the clear "Certainty Gap" where the imprecise group (dark blue) maintains higher confidence regardless of photo difficulty.
Critical Analysis & Conclusion
Takeaway
The study proves that human "errors" in precision are often a defense mechanism to maintain accuracy. By modeling this using belief functions, we can extract higher-quality "truth" from a crowd than by forcing them to be specific.
Limitations & Future Work
The current study uses a "simple" support mass function which only considers one focal element (the set of chosen answers). The authors suggest that future work should allow users to rank different sets of answers with different certainties—moving into consonant mass functions (possibility distributions). This would allow for even more nuanced human-AI knowledge transfer.
The Verdict: If you want better data from humans, let them be vague. Their certainty is worth more than their forced precision.
