LHNE: Synthesizing Structure and Content for Seamless User Identity Linkage

User identity linkage across social networks via linked heterogeneous network embedding

2018-04-23
Yaqing Wang, Chunyan Feng, Ling Chen, Hongzhi Yin, Caili Guo, Yunfei Chu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces LHNE (Linked Heterogeneous Network Embedding), a model for User Identity Linkage (UIL) across social networks. It bridges heterogeneous networks by jointly embedding structural (friendship links) and content (user-generated topics) information into a unified low-dimensional latent space, achieving state-of-the-art performance on Twitter-Flickr and DBLP datasets.

TL;DR

Connecting personal identities across different social platforms (e.g., linking a Twitter handle to a Flickr account) is a cornerstone of modern cross-network analytics. LHNE (Linked Heterogeneous Network Embedding) provides a breakthrough by refusing to treat social links and post content as separate entities. Instead, it embeds them into a unified "Interest-Friendship" manifold, resulting in a 47%+ performance boost over previous state-of-the-art models.

The "Independence" Fallacy in Prior Work

Most researchers previously assumed that "who you follow" (structure) and "what you say" (content) are independent variables. They would build one model for the social graph and another for text, then stitch them together like a Frankenstein’s monster.

However, the authors of LHNE identify a crucial Correlation Insight: A user following a celebrity is statistically likely to post about that celebrity. By failing to model this correlation in a shared space, previous methods lost critical information, especially for "isolated" users who lack social edges but are active content creators.

Methodology: The Unified Latent Space

LHNE addresses the heterogeneity problem through a multi-stage pipeline:

1. Topic-Based Denoising

Social media text is noisy (slang, ads, typos). LHNE uses Latent Dirichlet Allocation (LDA) to extract "Topics of Interest." This transforms raw text into a stable probability distribution over themes, effectively acting as a filter for irrelevant noise.

2. The Four Pillars of Linking

The model constructs a linked heterogeneous network consisting of four distinct sub-networks:

  • User-User Intra-network: Friendships within a platform.
  • User-Topic Intra-network: Interests within a platform.
  • User-User Inter-network: Transferring identity via known "anchor" users.
  • User-Topic Inter-network: Aligning topics across platforms (e.g., "Photography" on Flickr vs. "Camera" on Twitter).

LHNE Architecture

3. Joint Embedding Learning

The core of LHNE is its objective function. It doesn't just minimize the distance between friends; it minimizes the KL-divergence across all four sub-networks simultaneously. Using Negative Sampling, it ensures that "User A" from Twitter and "User A" from Flickr converge to the same point in a low-dimensional vector space.

Experimental Results: Dominance in Sparsity

The researchers tested LHNE against strong baselines like IONES and KNN on a Twitter-Flickr dataset.

Key Findings:

  • Strength in Isolation: In Flickr, roughly 40% of users had no friendship links. Standard structure-only models failed here. LHNE used their content "topics" as a proxy context to successfully identify them.
  • Efficiency: LHNE converges faster and requires fewer dimensions () to reach stability compared to its predecessors.

Performance Comparison Figure: LHNE showing superior Recall and Precision as the similarity threshold changes.

Critical Analysis: Why This Matters

The true value of LHNE lies in its robustness to data difficulty. In a world of increasing privacy protections and API limitations, we often cannot get a full social graph. LHNE proves that if you have even a sliver of content data, you can reconstruct the missing social context via the topic manifold.

Limitations: While powerful, the model relies on a "Seed Set" of known anchor users. In a purely cold-start environment where NO links are known, the inter-network transfer might struggle. Future work could look into zero-shot alignment using cross-lingual or cross-modal embeddings.

Conclusion

LHNE represents a shift from "Multi-view learning" (looking at different features separately) to "Unified Manifold Alignment." By recognizing that our interests and our friends are two sides of the same coin, it provides the most accurate bridge yet between our fragmented digital selves.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) instead of traditional network embedding for the User Identity Linkage task.
  • Which paper first introduced the concept of "Anchor Links" in heterogeneous social networks, and how does LHNE's transfer learning mechanism differ from that original work?
  • Explore how topic-based network embedding models like LHNE can be applied to cross-platform recommendation systems to solve the cold-start problem.
Contents
LHNE: Synthesizing Structure and Content for Seamless User Identity Linkage
1. TL;DR
2. The "Independence" Fallacy in Prior Work
3. Methodology: The Unified Latent Space
3.1. 1. Topic-Based Denoising
3.2. 2. The Four Pillars of Linking
3.3. 3. Joint Embedding Learning
4. Experimental Results: Dominance in Sparsity
5. Critical Analysis: Why This Matters
6. Conclusion