Active Data Augmentation: Harmonizing Crowdsourcing and Bootstrapping for Name Disambiguation

Bootstrapping active name disambiguation with crowdsourcing

2013-10-27
Yu Cheng, Zhengzhang Chen, Jiang Wang, Ankit Agrawal, Alok N. Choudhary
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Active Data Augmentation, a bootstrapping framework for author name disambiguation that integrates Active Learning with Crowdsourcing. By combining discriminative feature labeling and uncertainty-based sampling, the method achieves SOTA performance on DBLP and ArnetMiner datasets with significantly reduced human annotation effort.

TL;DR

The problem of "which 'John Smith' wrote this paper?" is a classic data science challenge known as name disambiguation. This paper presents a novel Active Data Augmentation framework that solves this by combining the "intuition" of discriminative features with the "judgment" of the crowd. By iteratively bootstrapping training data, the researchers achieved significant SOTA gains (+4.6% accuracy) while minimizing the need for expensive expert labeling.

The "Exploration-Exploitation" Dilemma in Disambiguation

Most modern name disambiguation systems rely on supervised learning. However, the bottleneck is always the same: Training Data.

Traditional Active Learning often falls into two traps:

  1. Sampling Bias: It hunts for the "hardest" samples (Uncertainty Sampling) but misses the bigger picture of the data distribution.
  2. Annotator Fatigue: Expecting experts to manually check thousands of pairs is unrealistic and error-prone.

The authors argue that we need a balance. We need Exploration (finding obvious, representative clusters using data features) and Exploitation (using human intelligence to solve the tricky, high-uncertainty edge cases).

Methodology: The Hybrid Bootstrapping Loop

The authors model the task as a Pairwise Graph Partitioning problem. Instead of just predicting a label, they model the conditional distribution of all match variables given observable features (like Co-authors, Titles, and Venues).

1. The Model Architecture

The framework utilizes a graph-based conditional model: This ensures that decisions aren't just local pairs but are globally consistent (e.g., if A=B and B=C, then A must equal C).

Overall Workflow Architecture

2. The Bootstrapping Algorithm

The "magic" happens in Algorithm 1:

  • Feature-Based Labeling (The Seed): The system automatically identifies pairs with extremely high similarity scores (e.g., identical co-authors and affiliations) to build an initial "pure" training set.
  • Crowdsourced Active Selection: The model finds the "most uncertain" pairs—where the probability of a match is near 0.5—and sends them to Amazon Mechanical Turk (MTurk).
  • Quality Control: To ensure the crowd isn't just guessing, each pair is checked by multiple workers, requiring consensus to be accepted.

Experimental Battleground: DBLP & ArnetMiner

The researchers tested their approach against four baselines, including Active Associative Sampling and standard Random Selection.

MethodAccuracy (DBLP)Macro-F1 (DBLP)
Random Selection0.7450.702
Active Selection0.7880.746
Active Data Augmentation (Ours)0.8680.831

Key Insights from Results:

  1. Initial Seeds Matter: Starting with just 30 feature-labeled samples significantly stabilizes the learning curve.
  2. The "S" Threshold: The number of discriminative samples () has a "sweet spot" (around 30-70). Too many high-confidence automatic labels can lead to overfitting, where the model ignores the nuances that only human labelers can catch.

Performance Comparison Graph Figure 1: The Macro-F1 score improves much faster per query compared to traditional active learning.

Critical Analysis & Conclusion

Why it works

The genius of this paper lies in Active Data Augmentation. By using discriminative features to handle the "obvious" cases, the query budget is reserved exclusively for "boundary" cases that truly define the classifier's margin.

Limitations

  • Crowd Reliability: While they used consensus, the paper doesn't deeply explore worker expertise variability.
  • Feature Dependency: The initial "pure clusters" depend heavily on the quality of metadata (like emails or affiliations), which might be missing in noisier datasets like social media.

Final Takeaway

This work demonstrates that the future of entity resolution isn't just "smarter models," but "smarter data acquisition." By treating crowdsourcing not just as a source of labels, but as a strategic component of a bootstrapping loop, we can resolve identity uncertainty at a fraction of the traditional cost.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Large Language Models (LLMs) as the 'oracle' in active learning frameworks for name disambiguation to replace or augment human crowdsourcing.
  • Which paper first proposed 'Conditional Models of Identity Uncertainty' for coreference tasks, and how does the pairwise graph partition in this paper extend that theory?
  • Explore how this bootstrapping active learning approach has been applied to cross-modal entity linking or social media user profiling beyond academic citation data.
Contents
Active Data Augmentation: Harmonizing Crowdsourcing and Bootstrapping for Name Disambiguation
1. TL;DR
2. The "Exploration-Exploitation" Dilemma in Disambiguation
3. Methodology: The Hybrid Bootstrapping Loop
3.1. 1. The Model Architecture
3.2. 2. The Bootstrapping Algorithm
4. Experimental Battleground: DBLP & ArnetMiner
4.1. Key Insights from Results:
5. Critical Analysis & Conclusion
5.1. Why it works
5.2. Limitations
5.3. Final Takeaway