Linking Shadows: High-Precision User Identification Across LinkedIn and Twitter

Similarity-Based User Identification Across Social Networks

2015-01-01
Katerina Zamani, Georgios Paliouras, Dimitrios Vogiatzis
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a similarity-based framework for identifying users across different social networks (specifically LinkedIn and Twitter) using a trainable combination of metrics. By aligning professional data, location, and social interactions, the authors achieve high-accuracy cross-platform linking (Recall up to 94.27%) to support information verification.

TL;DR

As our digital footprints scatter across platforms, the ability to "bridge" identities becomes crucial for verifying information. This paper introduces a supervised learning framework that combines professional achievements, geospatial data, and string similarity to identify the same individual across LinkedIn and Twitter with over 94% accuracy, effectively solving the name disambiguation problem in a professional context.

Problem & Motivation: The "John Smith" Paradox

In the context of the REVEAL project—aimed at verifying the trustworthiness of social media contributors—the biggest hurdle is Presence. If a journalist sees a claim on Twitter, how can they verify the author's credentials on a professional platform like LinkedIn?

The problem isn't just finding a name; it's disambiguation. A search for a common name returns dozens of profiles. Prior work often struggled with:

  • Data Sparsity: Users leave many profile fields (location, bio) blank.
  • Class Imbalance: In a set of 25 search results, only one (or zero) is a true match, making it easy for classifiers to "favor" the mismatch class.

Methodology: The Architecture of Similarity

The authors transform the identification task into a binary classification problem (Match vs. Mismatch). They generate a Similarity Vector based on five specialized dimensions.

User Attribute Alignment

1. The Five Pillar Metrics

  • Name (Jaro-Winkler): Handles slight variations in name spelling.
  • Description (Token Ratio): Compares short biographies by calculating the ratio of common keywords.
  • Location (Geospatial Semantic): Uses the GeoNames ontology. Instead of raw text matching, it calculates the ratio of "bounding box" areas. If someone says "Manhattan" on Twitter and "New York" on LinkedIn, the system recognizes the spatial subsumption.
  • Affiliation-Education (Smith-Waterman): Maps LinkedIn "Experience" to Twitter "User Mentions" (@tags), assuming professionals mention their employers in tweets.
  • Achievements (SoftTFIDF): A sophisticated metric that captures "similar" rather than identical tokens, linking job titles to bio descriptions (e.g., "Editor" vs. "Journalist").

2. The Hybrid Classifier: DTNB

The heart of the approach is the DTNB (Decision Table Naive Bayes) classifier. It combines the rule-based strengths of Decision Tables with the probabilistic robustness of Naive Bayes, assigned the maximum probability to determine a "Match" even under heavy data imbalance.

Experiments & Results: Professional Signals Win

The authors tested their approach on 262 "target users." A critical finding was the value of achievements. On LinkedIn, the "Achievements" metric alone achieved 87% recall, highlighting that professional context is more stable across platforms than names or locations.

ROC Curve Performance Fig 3: ROC curves showing DTNB outperforming Decision Trees and Naive Bayes in LinkedIn identification.

Key Quantitative Results:

TaskPrecisionRecallF-measure
LinkedIn Identification94.98%94.27%94.62%
Twitter Identification90.73%90.73%90.73%

The study also addressed missing values—a chronic issue in social data. They found that replacing missing scores with the average similarity score of other available fields outperformed setting a default constant (0.5) or using the median.

Critical Analysis & Conclusion

Takeaway

The research proves that cross-network identity is not just about the name; it’s about the semantic overlap of professional persona. By aligning unstructured text (mentions) with structured data (affiliations), we can bypass the noise of social media.

Limitations

  • Search Engine Dependency: The model relies on the top 25 results from native search engines. If the engine fails to retrieve the profile at all, the classifier can't find it.
  • Static Features: The model doesn't account for temporal changes in profiles (e.g., changing jobs).

Future Work

The authors suggest integrating SMOTE for better imbalance handling and exploring the identification of fake or compromised accounts—a vital next step in ensuring social media integrity.

Find Similar Papers

Try Our Examples

  • Find recent papers on cross-platform user identity linkage that specifically utilize Transformers or Large Language Models for embedding social media biographies.
  • What are the current SOTA methods for handling extreme class imbalance in entity resolution tasks beyond SMOTE and cost-sensitive learning?
  • Explore research that applies similarity-based user identification for detecting "sybil attacks" or fake accounts in decentralized social networks.
Contents
Linking Shadows: High-Precision User Identification Across LinkedIn and Twitter
1. TL;DR
2. Problem & Motivation: The "John Smith" Paradox
3. Methodology: The Architecture of Similarity
3.1. 1. The Five Pillar Metrics
3.2. 2. The Hybrid Classifier: DTNB
4. Experiments & Results: Professional Signals Win
4.1. Key Quantitative Results:
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work