Discovering Me Edges: Reconstructing User Identities Across the Social Internetworking Galaxy
Discovering Links among Social Networks
The paper introduces a novel recursive approach for discovering "me edges"—links connecting accounts of the same user across different social platforms. The method combines string similarity of usernames with a recursive common-neighbor analysis specifically designed for Social Internetworking Systems (SIS).
TL;DR
As users distribute their digital lives across platforms like Twitter, YouTube, and LinkedIn, a critical gap emerges: the missing "me edge"—the link that proves two accounts belong to the same person. This paper presents a recursive similarity framework that leverages usernames and neighborhood structures to "stitch" these fragmented identities together with over 85% accuracy, even when users don't explicitly link their accounts.
Problem & Motivation: The "Island" Problem of Social Networks
Most social network analysis treats individual platforms as isolated islands. While some users provide explicit "me" links (e.g., linking a Twitter profile to a personal blog), most do not.
The authors identify a major limitation in existing Link Prediction research: standard indices like the Jaccard or Salton Index rely on shared nodes. In a multi-network scenario, your friend "Alice" on Flickr has a different ID than "Alice" on Twitter. Because these IDs don't match, standard algorithms see zero common neighbors, resulting in a dismal sensitivity of ~1%.
Methodology: Recursive Similarity and Noise Reduction
The core innovation is a recursive similarity function that doesn't just look at names, but at the topology of relationships.
1. The Recursive Loop
The similarity between two users is a linear combination of their username similarity and the similarity of their best friend-pairs. Because those friends are also on different networks, the algorithm recursively calculates their similarity, effectively propagating identity confidence through the graph.
2. The V.I.P. Problem (Reduction Coefficient)
One major pitfall in social mining is the "Power User" effect. If two unrelated people both follow a global celebrity (like a V.I.P.), a naive algorithm might think they are the same person due to this "common" neighbor.
The authors introduce , a reduction coefficient that penalizes nodes with high degrees. This ensures that the similarity is driven by personal, niche connections rather than coincidental follows of public figures.

Experiments & Results: Precision vs. Recall
The team crawled a dataset of approximately 93,000 nodes across Twitter, LiveJournal, YouTube, and Flickr. They tested various string similarity metrics within their framework:
- QGrams: Proved to be the most "conservative," yielding the highest Precision (0.908).
- Needleman-Wunch: Proved the most "permissive," achieving a perfect Recall (1.000).
- State-of-the-Art Comparison: Traditional methods (Jaccard, Adamic-Adar) failed completely, while the proposed method maintained high performance across the board.

Critical Insight & Future Outlook
The value of this work lies in its Heuristic Intuition: it mirrors how a human might investigate a profile—"They have the same handle, and they both know a guy named Bob; let's see if those 'Bobs' are also the same person."
Limitations:
- Computational Complexity: While the authors prove a bound for visited nodes, a global-scale deployment on billions of users would require significant optimization.
- Privacy: The ability to link "separated" identities poses significant privacy risks, which the paper touches upon as a motivation for "completing profiles" but does not explore from an ethical standpoint.
Conclusion: This paper moves beyond simple text matching and treats Social Internetworking as a unified, albeit messy, graph. It provides a robust foundation for cross-platform behavioral analysis and enhanced crawling techniques.
