Discovering Me Edges: Reconstructing User Identities Across the Social Internetworking Galaxy

Discovering Links among Social Networks

2012-01-01
Francesco Buccafurri, Gianluca Lax, Antonino Nocera, Domenico Ursino
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel recursive approach for discovering "me edges"—links connecting accounts of the same user across different social platforms. The method combines string similarity of usernames with a recursive common-neighbor analysis specifically designed for Social Internetworking Systems (SIS).

TL;DR

As users distribute their digital lives across platforms like Twitter, YouTube, and LinkedIn, a critical gap emerges: the missing "me edge"—the link that proves two accounts belong to the same person. This paper presents a recursive similarity framework that leverages usernames and neighborhood structures to "stitch" these fragmented identities together with over 85% accuracy, even when users don't explicitly link their accounts.

Problem & Motivation: The "Island" Problem of Social Networks

Most social network analysis treats individual platforms as isolated islands. While some users provide explicit "me" links (e.g., linking a Twitter profile to a personal blog), most do not.

The authors identify a major limitation in existing Link Prediction research: standard indices like the Jaccard or Salton Index rely on shared nodes. In a multi-network scenario, your friend "Alice" on Flickr has a different ID than "Alice" on Twitter. Because these IDs don't match, standard algorithms see zero common neighbors, resulting in a dismal sensitivity of ~1%.

Methodology: Recursive Similarity and Noise Reduction

The core innovation is a recursive similarity function that doesn't just look at names, but at the topology of relationships.

1. The Recursive Loop

The similarity between two users is a linear combination of their username similarity and the similarity of their best friend-pairs. Because those friends are also on different networks, the algorithm recursively calculates their similarity, effectively propagating identity confidence through the graph.

2. The V.I.P. Problem (Reduction Coefficient)

One major pitfall in social mining is the "Power User" effect. If two unrelated people both follow a global celebrity (like a V.I.P.), a naive algorithm might think they are the same person due to this "common" neighbor.

The authors introduce , a reduction coefficient that penalizes nodes with high degrees. This ensures that the similarity is driven by personal, niche connections rather than coincidental follows of public figures.

Architecture/Formula Overview

Experiments & Results: Precision vs. Recall

The team crawled a dataset of approximately 93,000 nodes across Twitter, LiveJournal, YouTube, and Flickr. They tested various string similarity metrics within their framework:

  • QGrams: Proved to be the most "conservative," yielding the highest Precision (0.908).
  • Needleman-Wunch: Proved the most "permissive," achieving a perfect Recall (1.000).
  • State-of-the-Art Comparison: Traditional methods (Jaccard, Adamic-Adar) failed completely, while the proposed method maintained high performance across the board.

Performance Table

Critical Insight & Future Outlook

The value of this work lies in its Heuristic Intuition: it mirrors how a human might investigate a profile—"They have the same handle, and they both know a guy named Bob; let's see if those 'Bobs' are also the same person."

Limitations:

  • Computational Complexity: While the authors prove a bound for visited nodes, a global-scale deployment on billions of users would require significant optimization.
  • Privacy: The ability to link "separated" identities poses significant privacy risks, which the paper touches upon as a motivation for "completing profiles" but does not explore from an ethical standpoint.

Conclusion: This paper moves beyond simple text matching and treats Social Internetworking as a unified, albeit messy, graph. It provides a robust foundation for cross-platform behavioral analysis and enhanced crawling techniques.

Find Similar Papers

Try Our Examples

  • Find recent papers on cross-platform user identity linkage or anchor link prediction in heterogeneous social networks published after 2020.
  • Which paper first formally defined the "me edge" concept in the context of the Social Web, and how did it influence subsequent Social Internetworking analysis?
  • Analyze current research that applies Graph Neural Networks (GNNs) or embedding-based methods to solve the multi-network user alignment problem described in this study.
Contents
Discovering Me Edges: Reconstructing User Identities Across the Social Internetworking Galaxy
1. TL;DR
2. Problem & Motivation: The "Island" Problem of Social Networks
3. Methodology: Recursive Similarity and Noise Reduction
3.1. 1. The Recursive Loop
3.2. 2. The V.I.P. Problem (Reduction Coefficient)
4. Experiments & Results: Precision vs. Recall
5. Critical Insight & Future Outlook