Corleone: Scaling Entity Matching via Hands-Off Crowdsourcing

Corleone: Hands-off crowdsourcing for entity matching

2014-01-01
Chaitanya Gokhale, Sanjib Das, Anhai Doan, Jeffrey F. Naughton, Narasimhan Rampalli, Jude Shavlik, Xiaojin Zhu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Corleone, the first Hands-Off Crowdsourcing (HOC) framework for Entity Matching (EM). It automates the entire EM workflow—including blocking, matching, accuracy estimation, and iterative refinement—using an active learning-based Random Forest and crowd workers, eliminating the need for expert developers.

TL;DR

Entity Matching (EM)—the art of identifying different records referring to the same real-world entity—has long been a bottleneck in data science. While crowdsourcing helped, it still required a developer to "babysit" the process. Corleone breaks this barrier by introducing Hands-Off Crowdsourcing (HOC), a system that uses the crowd for every step of the workflow, achieving SOTA accuracy (up to 96.5% F1) without a single line of code from the user.

The "Developer Bottleneck" in Crowdsourcing

Before Corleone, crowdsourced EM followed a "hybrid" model:

  1. Developer: Writes complex pearl/python scripts for blocking (filtering out obvious non-matches).
  2. Crowd: Labels a few remaining ambiguous pairs.
  3. Developer: Evaluates results, tunes the model, and iterates.

For a large enterprise like Amazon or Walmart with thousands of categories, hiring developers for every single EM task is a scaling nightmare. For a journalist or small business owner, it’s an impossible barrier to entry.

Methodology: How Corleone Replaces the Expert

The core innovation of Corleone is its ability to turn "dumb" crowd clicks into "smart" machine-readable rules.

1. Crowdsourced Blocking

How do you get a crowd worker to write a rule like if price_diff > 20 then no_match? You don't. Corleone samples a small portion of the data, uses active learning to build a Random Forest, and then extracts rules from the tree branches. The crowd simply labels pairs, and the system deduces the logic.

Corleone Architecture

2. Active Learning with a "Smart" Stop Button

Training a matcher usually requires thousands of labels. Corleone uses Active Learning to pick the most informative pairs (those where the Random Forest models disagree most).

  • The Problem: Crowds are noisy. If you keep training with noisy labels, model accuracy eventually tanks.
  • The Solution: A confidence-based stopping mechanism. Corleone monitors a validation set and stops training once the internal "agreement" of the Random Forest peaks, preventing over-exposure to crowd errors.

3. Solving the Skewed Data Problem

In EM, matches are like needles in a haystack (often < 1% of the data). Standard precision/recall estimation fails because a random sample rarely finds enough "needles." Corleone uses its learned rules to "reduce" the haystack, increasing the density of positives so that accuracy can be estimated with 90% fewer labels compared to traditional methods.

Experiments & Results

Corleone was tested against real-world datasets: Restaurants, Citations, and Electronics Products.

  • Accuracy: On the "Products" dataset (the hardest one), Corleone achieved an 89.3% F1, smashing the baseline of 69.5%.
  • Efficiency: Despite doing everything automatically, the cost was surprisingly low—just 256 for datasets with millions of potential combinations.

Experimental Results Comparison

Critical Insight: Why This Matters

The takeaway is profound: Crowdsourcing isn't just for labeling data; it's for cleaning and generating models. Corleone proves that an end-to-end HOC system can be "The Godfather" of its domain—managing the complex "mob" of crowd workers in a hands-off fashion to deliver professional-grade results.

Limitations & Future Work

  • Crowd Sensitivity: While robust to some noise, a 20% error rate in crowd labels still causes performance drops.
  • Cloud Scaling: Future iterations need better Hadoop/Spark integration to handle billions of records in hours rather than days.

Conclusion

Corleone represents a paradigm shift. By removing the developer from the loop, it democratizes high-performance Entity Matching for everyone from Fortune 500 enterprises to "data enthusiasts."

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend the concept of "Hands-Off Crowdsourcing" (HOC) to high-level data integration tasks like schema matching or information extraction.
  • Identify the foundational research on using Active Learning with Random Forests for entity matching that Corleone's rule-extraction methodology is built upon.
  • Explore how subsequent studies have improved upon Corleone's confidence-based stopping criteria to better handle noise in crowd-labeled data.
Contents
Corleone: Scaling Entity Matching via Hands-Off Crowdsourcing
1. TL;DR
2. The "Developer Bottleneck" in Crowdsourcing
3. Methodology: How Corleone Replaces the Expert
3.1. 1. Crowdsourced Blocking
3.2. 2. Active Learning with a "Smart" Stop Button
3.3. 3. Solving the Skewed Data Problem
4. Experiments & Results
5. Critical Insight: Why This Matters
5.1. Limitations & Future Work
6. Conclusion