Beyond Names: Using Social "Fingerprints" to Solve Identity Ambiguity on the Web
Social Relationships as a Means for Identifying an Individual in Large Information Spaces
The paper introduces a novel person-identification method for large information spaces like the Web and DBLP. It leverages Social Network analysis, combining Levenshtein distance for syntactic matching and a modified "Connected Triple" metric for semantic relationship comparison to disambiguate individuals with similar names.
TL;DR
Identified a person in a sea of "John Smiths" is a classic data science challenge. This paper argues that who you know is more unique than what your name is. By comparing the social network topology extracted from web pages against background knowledge using a refined Connected Triple metric, the authors provide a robust framework for name disambiguation that outperforms traditional heuristic-based models.
The Identity Crisis in Information Spaces
When you search for a researcher on DBLP or a professional on the Web, you often encounter two types of noise:
- Multi-referent ambiguity: Five different "Michael Smiths" appearing in the same search result.
- Multi-morphic ambiguity: "M. Smith," "Mike Smith," and "Michael Smith" all referring to the same person.
While most systems try to solve this using keywords or rare personal attributes (like birth dates, which are often unavailable due to privacy), this paper shifts the focus to Social Networks. The core intuition is powerful: while two people might share a name, the probability of them sharing the same set of professional or personal connections is virtually zero.
Methodology: Extracting and Comparing Social "Graphs"
The authors propose a dual-phase pipeline:
1. Social Network Extraction
The system crawls web pages starting from a target URL. It extracts names using an English dictionary and identifies relationships based on the link structure between pages. If Page A links to Page B, and names appear on both, a relationship edge is formed.
2. The Semantic Bridge: Modified Connected Triples
Comparing two graphs is computationally expensive. To bridge this, the authors use:
- Syntactic Filter: Levenshtein distance ensures we only compare nodes that look similar (e.g., "Ivan" and "Ivo").
- Semantic Comparison: This is where the magic happens. The authors use a Connected Triple—a subgraph of three vertices where two people (the candidates for identification) are both connected to a common third party, but not necessarily to each other.
The Innovation: Local vs. Global Focus The original Connected Triple formula normalized similarity against the maximum number of triples in the entire graph. The authors realized this "dilutes" the similarity score in large datasets. They modified the formula to normalize only against triples where at least one of the compared persons is involved:
The modified formula focuses on the "local" social neighborhood, making the similarity score much more sensitive to actual identity matches.
Experiments and Results
The authors tested their method on the Slovak Companies Register (a massive dataset of 300k nodes) and university staff pages.
- Performance vs. Baseline: The method identified significantly more duplicates than standard domain-specific heuristics (e.g., matching names + addresses).
- Modified vs. Original: The local-normalization approach (Formula 2) showed a massive leap in recall—identifying duplicates that the original metric missed entirely in the "Novak" and "Havran" datasets.
Table showing the superiority of the modified method in recall and precision compared to original metrics.
Critical Insight: The Strength of Weak Links
The most profound takeaway is that identity is a relational property. Instead of looking at the entity, we should look around it.
Limitations
- Dictionary Dependence: The current extraction tool struggles with non-English names (recall dropped on foreign datasets).
- Weighting: The model treats all social links as equal. In reality, a co-author link is much "stronger" evidence of identity than a mere mention on the same news page.
Future Outlook
This approach is highly applicable to Social Discovery. Imagine a social portal (like LinkedIn) that doesn't just match you by name but crawls your personal website to suggest friends based on the "web environment" of your social links. By moving from string matching to topology matching, we move closer to a truly "Semantic Web."
