Bridging the Silos: Iterative Profile Matching Across Social Networks

Matching User Profiles Across Social Networks

2014-01-01
Nacéra Bennacer, Coriane Nana Jipmo, Antonio Penta, Gianluca Quercini
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents an iterative algorithm for matching user profiles across multiple social networks by combining network topology and publicly available attributes. Specifically, it uses existing "cross-links" to identify candidate pairs among friends and then applies a hierarchical set of matching rules, achieving a precision of 94% on a dataset of 2 million profiles.

TL;DR

Researchers have developed a highly accurate (94% precision) iterative algorithm to link the same individual's profiles across different social networks. By combining the "who you know" (topology) with "what you share" (attributes) and treating the problem as an expanding search, they successfully mapped identities across a massive dataset of 2 million profiles.

Background & Motivation: The Fragmented Digital Self

In the Web 2.0 era, your digital identity is scattered. You might be a professional on LinkedIn, a photographer on Flickr, and a micro-blogger on Twitter. While these platforms are silos, the social connections between them are not.

The core challenge—User Identity Linkage (UIL)—is difficult because:

  1. Data Inconsistency: Users provide different information on different sites.
  2. Scale: Comparing every pair of users across two platforms like Flickr and Twitter is computationally impossible ().
  3. Privacy and Noise: Users often use fake names or omit emails, making simple string matching unreliable.

Methodology: The Power of Iteration

The authors suggest that we shouldn't just look for "Who is Bob?" but rather "Who does Bob's friend know?".

1. Candidate Selection (Pruning the Search Space)

Instead of comparing everyone to everyone, the algorithm uses existing cross-links as seeds. If Alice and Bob are friends on Flickr, and we already know Alice's Twitter account, Bob’s Twitter account is likely to be found among Alice’s Twitter friends.

2. Hierarchical Matching Rules

Once candidates are narrowed down, the system applies rules of varying "orders" ():

  • High Confidence: Comparison of unique identifiers like email or links to the same personal website.
  • Medium Confidence: Sophisticated string distances (Levenshtein for usernames, Jaccard for names).
  • Combined Proof: If both the username and the name are somewhat similar, the confidence increases.

Algorithm Workflow Note: The algorithm iteratively discovers new matches (D), which then serve as seeds for the next round of candidate selection.

Experiments: Scaling to 2 Million Nodes

The researchers didn't just test this on a small lab sample. They scaled it to a "social internetwork" involving Flickr, LiveJournal, Twitter, and YouTube.

NetworkNodesFriendship Links
Flickr1,814,40515,415,083
LiveJournal211,0452,093,737
Total~2,000,000~17,500,000

Key Findings:

  • Precision: 94% of the discovered links were manually verified as correct.
  • Iterative Gains: The algorithm doesn't stop after the first pass. In their test, it took four iterations to reach convergence, discovering hundreds of new links in the 2nd and 3rd rounds that were invisible at the start.
  • Attribute Weight: "Usernames" and "Links to other profiles" are the strongest predictors. Interestingly, "Real Names" are often too ambiguous or fake to be used alone.

Experimental Results

Critical Analysis: Why This Matters

The brilliance of this work lies in its simplicity and scalability. By using topology to prune the search space, it avoids the "big data" trap of comparing trillions of pairs.

Limitations:

  • Dependency on Seeds: The algorithm needs an initial set of users who have manually linked their accounts (e.g., "Find me on Twitter" links in a Flickr bio).
  • Privacy Concerns: As these algorithms become more accurate, the ability to remain "anonymous" across different facets of one's life diminishes, raising significant ethical questions about data aggregation.

Conclusion

This research demonstrates that your social circle is a powerful identifier. Even if you change your name or handle, the structure of your relationships acts as a "social fingerprint" that can be used to reconnect your digital fragments with surgical precision.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Graph Neural Networks (GNNs) or embedding-based methods to solve the User Identity Linkage (UIL) problem across social networks.
  • Which paper originally proposed the "self-disclosure of cross-links" as a ground truth for social network alignment, and how has the reliability of this anchor changed with modern privacy settings?
  • Explore how iterative profile matching algorithms have been adapted to handle anonymized or encrypted social graphs where attribute information is entirely missing.
Contents
Bridging the Silos: Iterative Profile Matching Across Social Networks
1. TL;DR
2. Background & Motivation: The Fragmented Digital Self
3. Methodology: The Power of Iteration
3.1. 1. Candidate Selection (Pruning the Search Space)
3.2. 2. Hierarchical Matching Rules
4. Experiments: Scaling to 2 Million Nodes
4.1. Key Findings:
5. Critical Analysis: Why This Matters
6. Conclusion