AuthCrowd: Resolving Academic Identity Crisis through Professional Crowdsourcing

AuthCrowd: Author Name Disambiguation and Entity Matching using Crowdsourcing

2021-05-05
António Correia, Diogo Guimarães, Dennis Paulino, Shoaib Jameel, Daniel Schneider, Benjamim Fonseca, Hugo Paredes
Summary
Problem
Method
Results
Takeaways
Abstract

AuthCrowd is a crowdsourcing system designed for Author Name Disambiguation (AND) and entity matching in bibliometric datasets. It integrates automated data acquisition with a human-in-the-loop framework, achieving over 80% accuracy in complex metadata transcription tasks.

TL;DR

The exponential growth of scientific literature has made Author Name Disambiguation (AND) a nightmare for digital libraries. AuthCrowd introduces a specialized crowdsourcing framework that leverages human intuition to solve the "homonym/synonym" problem in bibliometrics. By decomposing complex matching into digestible micro-tasks, the system achieves over 80% accuracy in metadata validation, proving that the human-in-the-loop approach is vital for high-quality scientometric data.

The "Identity Crisis" in Scientometrics

In the world of big data, "John Smith" is more than a name; it is a point of failure. Modern algorithms struggle when:

  • An author changes affiliations (Temporal Ambiguity).
  • Multiple authors share near-identical names (Homonymy).
  • Names are misspelled or inconsistently initialized (Synonymy).

While Unsupervised Learning (Clustering) and Network Embeddings are the current SOTA, they are brittle when faced with missing data. The authors of AuthCrowd argue that we have hit a plateau with pure ML and must return to Human Intelligence—specifically, a structured, scalable "Crowd" approach—to clean the metadata registry.

Methodology: The AuthCrowd Architecture

The system treats disambiguation not as a single calculation, but as a multi-stage pipeline:

  1. Task Design: Presenting publications in varied visual formats (e.g., Name and Affiliation in Same Row (NAR) vs. Separated (NAS)) to minimize cognitive load.
  2. Validation: Using "Ground Truth" questions to filter out low-quality contributors.
  3. Aggregation: Employing techniques like Majority Voting and HoneyPot to distill a single, high-confidence answer from diverse crowd inputs.

AuthCrowd Task Interface Fig 1: The AuthCrowd interface demonstrating side-by-side article comparison for entity matching.

Experimental Insights: Does the Crowd Deliver?

The study evaluated parameters like Visual Representation and Confidence Scoring.

Key Findings:

  • The Difficulty Gradient: Tasks were categorized from A (Hard - different affiliations) to C (Easy). The crowd performed best when the task involved identifying a "Corresponding Author," reaching SOTA-level precision.
  • Behavioral Efficiency: Log analysis showed a "learning curve." As users became familiar with the interface, the "Average Number of Clicks" decreased while accuracy remained stable or improved.
  • The Power of Confidence: There was a strong correlation between a participant's self-reported "Confidence Level" and the actual "Correctness" of the result. High-confidence positive answers (Scale 4-5) yielded over 75% accuracy.

Task Overview and Difficulty Table 1: Overview of task complexity vs. accuracy outcomes.

Why This Matters: Moving Beyond the Algorithm

The core contribution of AuthCrowd is the demonstration that crowdsourcing is a "Reasonable Method" for bibliometric IR. Purely algorithmic approaches often treat data points as static; AuthCrowd treats them as part of a narrative that human researchers can navigate more effectively—such as noticing that a change in affiliation is consistent with a typical career progression.

Limitations & Future Work

  • Scalability: While 24 participants provided a proof-of-concept, scaling to millions of papers requires a "Hybrid Crowd-Algorithm" strategy where ML handles the 90% easy cases, and AuthCrowd handles the 10% high-entropy edge cases.
  • Interface Overlap: User feedback suggested that side-by-side zoom features occasionally caused text overlap, suggesting a need for more responsive UI design.

Conclusion

AuthCrowd isn't just a tool; it's a paradigm shift. It suggests that the future of authoritative scientific databases depends on a Collaborative Intelligence model, where the expert crowd acts as the ultimate supervisor for the AI.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize active learning combined with crowdsourcing for large-scale author name disambiguation tasks.
  • What are the inherent theoretical differences between graph-based unsupervised clustering and human-in-the-loop verification in entity resolution, as discussed in foundational scientometrics literature?
  • Explore how crowdsourcing frameworks like AuthCrowd can be extended to resolve cross-modal entity matching in academic datasets involving both text and images (e.g., figure-author attribution).
Contents
AuthCrowd: Resolving Academic Identity Crisis through Professional Crowdsourcing
1. TL;DR
2. The "Identity Crisis" in Scientometrics
3. Methodology: The AuthCrowd Architecture
4. Experimental Insights: Does the Crowd Deliver?
4.1. Key Findings:
5. Why This Matters: Moving Beyond the Algorithm
5.1. Limitations & Future Work
6. Conclusion