Mining the Past: How Social Network Properties Reveal Name Inconsistencies in DBLP
Learning from the Past: An Analysis of Person Name Corrections in DBLP Collection and Social Network Properties of Affected Entities
This paper presents a longitudinal study of name-related data quality in the DBLP digital library by mining person name corrections over a ten-year period. It leverages a "historic DBLP collection" to categorize errors—such as synonyms and homonyms—and analyzes their distribution within dynamic social networks of coauthorship and publication streams.
TL;DR
Researchers at the University of Trier have analyzed ten years of historical data from the DBLP bibliographic project to understand why name errors occur. By tracking how "defective" names were corrected over time, they discovered that an author's position in a social network—who they coauthor with and where they publish—is a strong predictor of whether their name record is accurate or requires correction.
Background: The Identity Crisis in Digital Libraries
Digital libraries like DBLP are the backbone of academic research, but they face a persistent challenge: Name Ambiguity. This manifests in two primary ways:
- Homonyms: Multiple people sharing the same name (e.g., the 19 "Wei Wangs" in DBLP).
- Synonyms: One person being listed under multiple names due to spelling errors, transcription differences, or name changes.
While automated algorithms exist, they are rarely perfect. The gold standard remains manual correction, often triggered by "community feedback" from authors themselves. This paper asks: Can we learn from these past human-made corrections to predict future errors?
Methodology: The Historic DBLP Framework
The authors reconstructed the history of DBLP by mining 3,300 backups from 1999 to 2009. They defined four types of modifications to track how data quality improved over time:
- Rename: A simple label change (e.g., "H. Schweppe" to "Heinz Schweppe").
- Merge: Combining two names into one (resolving synonyms).
- Split: Dividing one name into two (resolving homonyms).
- Distribute: Reassigning a specific paper from one author to another.
Analyzing Network DNA
The research evaluated these entities through two distinct social networks:
- Collaboration Network (C): A graph where nodes are authors and edges represent coauthorship.
- Name-Stream Network (S): A two-mode network linking authors to the "streams" (conferences or journals) where they publish.
Fig 1: Conceptual mapping between real-world persons and DBLP name entities.
Key Insights: Error Profiles
The study found significant differences in how "incorrect" names exist within the network:
- Local Connectivity: Names that were eventually split or merged tended to have a very low Clustering Coefficient. Essentially, if your "coauthors" don't know each other (not forming a clique), you might actually be two different people mistakenly merged into one record.
- Global Centrality: Interestingly, defective entities often showed high Betweenness Centrality. This suggests that prominent authors with high publication counts are more likely to have errors—partly because they have more opportunities for data entry mistakes, and partly because the community monitors their records more closely.
Fig 2: Distribution of node degrees in the collaboration network, showing how split/distribute candidates often have larger neighborhoods than stable nodes.
A New Heuristic: Assessing "Reliable Areas"
The authors proposed that errors aren't just random—they cluster. Some publishers or conferences (streams) have "noisier" metadata than others. By calculating a Stream Node Risk (rS) and a Collaboration Risk (rC), they can assign a reliability score to any given name.
In their evaluation:
- Safe Zones: Name entities in the top-tier "Reliable Areas" had very low correction rates.
- Danger Zones: Entities in the bottom-tier "Error Prone Areas" (often containing Spanish or Chinese names with incomplete first names) were 3x more likely to be corrected in the following two years.
Critical Analysis & Future Outlook
This work provides a unique "Ground Truth" for the research community. While most papers try to solve disambiguation with the current state of a database, this paper proves that the evolutionary history of the database is a goldmine for improving data quality.
Limitations: The study is biased toward errors that DBLP’s current tools and community are good at catching. Errors that remain undiscovered by humans are, by definition, missing from this analysis.
Takeaway: For builders of knowledge graphs and digital libraries, the message is clear: don't just look at the node—look at the neighborhood. If an author publishes in streams with high historical error rates and has a "star-shaped" coauthor network, their record likely requires a closer look.
