Active Data Augmentation: Harmonizing Crowdsourcing and Bootstrapping for Name Disambiguation
Bootstrapping active name disambiguation with crowdsourcing
The paper introduces Active Data Augmentation, a bootstrapping framework for author name disambiguation that integrates Active Learning with Crowdsourcing. By combining discriminative feature labeling and uncertainty-based sampling, the method achieves SOTA performance on DBLP and ArnetMiner datasets with significantly reduced human annotation effort.
TL;DR
The problem of "which 'John Smith' wrote this paper?" is a classic data science challenge known as name disambiguation. This paper presents a novel Active Data Augmentation framework that solves this by combining the "intuition" of discriminative features with the "judgment" of the crowd. By iteratively bootstrapping training data, the researchers achieved significant SOTA gains (+4.6% accuracy) while minimizing the need for expensive expert labeling.
The "Exploration-Exploitation" Dilemma in Disambiguation
Most modern name disambiguation systems rely on supervised learning. However, the bottleneck is always the same: Training Data.
Traditional Active Learning often falls into two traps:
- Sampling Bias: It hunts for the "hardest" samples (Uncertainty Sampling) but misses the bigger picture of the data distribution.
- Annotator Fatigue: Expecting experts to manually check thousands of pairs is unrealistic and error-prone.
The authors argue that we need a balance. We need Exploration (finding obvious, representative clusters using data features) and Exploitation (using human intelligence to solve the tricky, high-uncertainty edge cases).
Methodology: The Hybrid Bootstrapping Loop
The authors model the task as a Pairwise Graph Partitioning problem. Instead of just predicting a label, they model the conditional distribution of all match variables given observable features (like Co-authors, Titles, and Venues).
1. The Model Architecture
The framework utilizes a graph-based conditional model: This ensures that decisions aren't just local pairs but are globally consistent (e.g., if A=B and B=C, then A must equal C).

2. The Bootstrapping Algorithm
The "magic" happens in Algorithm 1:
- Feature-Based Labeling (The Seed): The system automatically identifies pairs with extremely high similarity scores (e.g., identical co-authors and affiliations) to build an initial "pure" training set.
- Crowdsourced Active Selection: The model finds the "most uncertain" pairs—where the probability of a match is near 0.5—and sends them to Amazon Mechanical Turk (MTurk).
- Quality Control: To ensure the crowd isn't just guessing, each pair is checked by multiple workers, requiring consensus to be accepted.
Experimental Battleground: DBLP & ArnetMiner
The researchers tested their approach against four baselines, including Active Associative Sampling and standard Random Selection.
| Method | Accuracy (DBLP) | Macro-F1 (DBLP) |
|---|---|---|
| Random Selection | 0.745 | 0.702 |
| Active Selection | 0.788 | 0.746 |
| Active Data Augmentation (Ours) | 0.868 | 0.831 |
Key Insights from Results:
- Initial Seeds Matter: Starting with just 30 feature-labeled samples significantly stabilizes the learning curve.
- The "S" Threshold: The number of discriminative samples () has a "sweet spot" (around 30-70). Too many high-confidence automatic labels can lead to overfitting, where the model ignores the nuances that only human labelers can catch.
Figure 1: The Macro-F1 score improves much faster per query compared to traditional active learning.
Critical Analysis & Conclusion
Why it works
The genius of this paper lies in Active Data Augmentation. By using discriminative features to handle the "obvious" cases, the query budget is reserved exclusively for "boundary" cases that truly define the classifier's margin.
Limitations
- Crowd Reliability: While they used consensus, the paper doesn't deeply explore worker expertise variability.
- Feature Dependency: The initial "pure clusters" depend heavily on the quality of metadata (like emails or affiliations), which might be missing in noisier datasets like social media.
Final Takeaway
This work demonstrates that the future of entity resolution isn't just "smarter models," but "smarter data acquisition." By treating crowdsourcing not just as a source of labels, but as a strategic component of a bootstrapping loop, we can resolve identity uncertainty at a fraction of the traditional cost.
