CrowdEC: Reducing Costs and Redundancy in Crowdsourced Entity Collection
Incentive-Based Entity Collection Using Crowdsourcing
This paper introduces CrowdEC, an incentive-based crowdsourced entity collection framework designed to gather complete and correct entity sets (e.g., NBA players, universities) while minimizing costs. It achieves state-of-the-art performance by combining a worker elimination strategy with a dynamic bonus-based incentive pricing mechanism.
Executive Summary
CrowdEC is a sophisticated framework designed to solve the "popularity bias" in crowdsourcing, where workers tend to provide the same easy-to-recall entities (e.g., Lebron James) while ignoring the long-tail (e.g., bench players). By integrating a Worker Elimination module and an Incentive Pricing mechanism, the framework ensures that we only pay for high-quality, distinct data. This work represents a significant leap from passive statistical estimation to active participant management in the crowdsourcing ecosystem.
The Pain Point: The Expensive Duplication Problem
Crowdsourcing is the go-to method for building knowledge bases. However, it faces two structural flaws:
- Wasteful Redundancy: If you ask 100 workers for an NBA player, 90 might say "Stephen Curry." You've paid for 100 tasks but gained almost no new knowledge.
- The Quality Vacuum: In open-ended tasks ("Give me a university name"), there is no pre-defined "Golden Set" to test worker honesty, leading to "cheating" or low-effort submissions.
Methodology: The Architecture of Incentives
CrowdEC shifts the paradigm by treating entity collection as a dynamic optimization problem.
1. Worker Elimination (Utility Maximization)
Not all workers are equal. CrowdEC calculates a Worker Utility score based on:
- Throughput ( / ): How many unique entities a worker provides per request.
- Error Bounding (): Using a Bayesian approach to estimate the probability that a submitted entity actually belongs to the target domain ().

2. Incentive Pricing (The Bonus Strategy)
Since platforms like Amazon Mechanical Turk (AMT) don't support dynamic price negotiation, CrowdEC uses a Bonus-based retry logic:
- NoBonus: Standard task, base reward.
- Bonus: If the worker provides a duplicate, the system notifies them and offers a bonus for a distinct new entry. This turns a "failed" attempt into a successful data point at a lower marginal cost than a new task.

Experiments and Results
The authors validated the system using datasets like ActiveNBA (450 entities) and TopUniv (100 entities).
- Cost Efficiency: CrowdEC significantly flattened the cost curve. By the time 420 NBA players were collected, CrowdEC's costs were nearly 3x lower than traditional "Enumeration" methods.
- Precision/Recall Balance: Unlike "No-Selection" baselines where precision drops as workers get tired or bored, CrowdEC’s quality control kept precision above 95%.

Critical Analysis & Takeaways
The brilliance of CrowdEC lies in its Online Algorithm. It makes pricing decisions in real-time without needing to know the future distribution of worker requests, achieving an approximation ratio close to the theoretical optimum (Theorem 2).
Future Outlook: While highly effective, CrowdEC assumes that "entity resolution" (deciding if "S. Curry" and "Stephen Curry" are the same) is handled externally. Integrating a real-time, human-in-the-loop entity resolution module could further refine the "Distinctness" metric and provide even higher cost savings for massive datasets like the 4,260-item "AllNBA" list.
Conclusion: CrowdEC proves that financial incentives, when applied algorithmically based on data distinctness rather than just participation, can solve one of the oldest problems in human-powered data collection.
