Beyond Names: Using Social "Fingerprints" to Solve Identity Ambiguity on the Web

Social Relationships as a Means for Identifying an Individual in Large Information Spaces

2010-01-01
Katarína Kostková, Michal Barla, Mária Bieliková
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel person-identification method for large information spaces like the Web and DBLP. It leverages Social Network analysis, combining Levenshtein distance for syntactic matching and a modified "Connected Triple" metric for semantic relationship comparison to disambiguate individuals with similar names.

TL;DR

Identified a person in a sea of "John Smiths" is a classic data science challenge. This paper argues that who you know is more unique than what your name is. By comparing the social network topology extracted from web pages against background knowledge using a refined Connected Triple metric, the authors provide a robust framework for name disambiguation that outperforms traditional heuristic-based models.

The Identity Crisis in Information Spaces

When you search for a researcher on DBLP or a professional on the Web, you often encounter two types of noise:

  1. Multi-referent ambiguity: Five different "Michael Smiths" appearing in the same search result.
  2. Multi-morphic ambiguity: "M. Smith," "Mike Smith," and "Michael Smith" all referring to the same person.

While most systems try to solve this using keywords or rare personal attributes (like birth dates, which are often unavailable due to privacy), this paper shifts the focus to Social Networks. The core intuition is powerful: while two people might share a name, the probability of them sharing the same set of professional or personal connections is virtually zero.

Methodology: Extracting and Comparing Social "Graphs"

The authors propose a dual-phase pipeline:

1. Social Network Extraction

The system crawls web pages starting from a target URL. It extracts names using an English dictionary and identifies relationships based on the link structure between pages. If Page A links to Page B, and names appear on both, a relationship edge is formed.

2. The Semantic Bridge: Modified Connected Triples

Comparing two graphs is computationally expensive. To bridge this, the authors use:

  • Syntactic Filter: Levenshtein distance ensures we only compare nodes that look similar (e.g., "Ivan" and "Ivo").
  • Semantic Comparison: This is where the magic happens. The authors use a Connected Triple—a subgraph of three vertices where two people (the candidates for identification) are both connected to a common third party, but not necessarily to each other.

The Innovation: Local vs. Global Focus The original Connected Triple formula normalized similarity against the maximum number of triples in the entire graph. The authors realized this "dilutes" the similarity score in large datasets. They modified the formula to normalize only against triples where at least one of the compared persons is involved:

Formula 2 The modified formula focuses on the "local" social neighborhood, making the similarity score much more sensitive to actual identity matches.

Experiments and Results

The authors tested their method on the Slovak Companies Register (a massive dataset of 300k nodes) and university staff pages.

  • Performance vs. Baseline: The method identified significantly more duplicates than standard domain-specific heuristics (e.g., matching names + addresses).
  • Modified vs. Original: The local-normalization approach (Formula 2) showed a massive leap in recall—identifying duplicates that the original metric missed entirely in the "Novak" and "Havran" datasets.

Comparison Chart Table showing the superiority of the modified method in recall and precision compared to original metrics.

Critical Insight: The Strength of Weak Links

The most profound takeaway is that identity is a relational property. Instead of looking at the entity, we should look around it.

Limitations

  • Dictionary Dependence: The current extraction tool struggles with non-English names (recall dropped on foreign datasets).
  • Weighting: The model treats all social links as equal. In reality, a co-author link is much "stronger" evidence of identity than a mere mention on the same news page.

Future Outlook

This approach is highly applicable to Social Discovery. Imagine a social portal (like LinkedIn) that doesn't just match you by name but crawls your personal website to suggest friends based on the "web environment" of your social links. By moving from string matching to topology matching, we move closer to a truly "Semantic Web."

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) or Graph Embedding techniques for the specific task of name disambiguation in professional databases like DBLP or LinkedIn.
  • What is the original definition of the "Connected Triple" metric as proposed by Reuther et al. (2006), and how have other researchers modified it for entity resolution?
  • Explore current SOTA methods for cross-platform identity linking (e.g., matching a GitHub profile to a Twitter account) that utilize social graph topology rather than just profile metadata.
Contents
Beyond Names: Using Social "Fingerprints" to Solve Identity Ambiguity on the Web
1. TL;DR
2. The Identity Crisis in Information Spaces
3. Methodology: Extracting and Comparing Social "Graphs"
3.1. 1. Social Network Extraction
3.2. 2. The Semantic Bridge: Modified Connected Triples
4. Experiments and Results
5. Critical Insight: The Strength of Weak Links
5.1. Limitations
6. Future Outlook