SCAN-PS: Defeating the Cold-Start Demon in Social Networks through Cross-Platform Alignment
Predicting Social Links for New Users across Aligned Heterogeneous Social Networks
The paper introduces SCAN-PS, a supervised link prediction framework designed specifically for new users by leveraging multiple aligned heterogeneous social networks. It utilizes "anchor links" to transfer knowledge from source networks to a target network, achieving SOTA performance in cold-start scenarios.
TL;DR
Predicting who a new user will follow is notoriously difficult because they have no history. SCAN-PS solves this by "borrowing" the user's history from other platforms (like using a user's Twitter activity to suggest friends on Foursquare). By combining Personalized Sampling (to bridge the gap between old and new users) and Cross-Network Feature Fusion, it achieves high accuracy even for "brand new" accounts with zero local data.
The "New User" Paradox
In social network science, the "first impression" is everything. If a platform recommends relevant connections immediately, a new user stays; if not, they churn. However, traditional link prediction models (like Common Neighbors or Jaccard Coefficient) rely on existing connections—which new users don't have.
The authors identify two fatal flaws in prior work:
- The Distribution Gap: Old users have dense metadata and many links; new users are sparse. Training on one and testing on the other violates the i.i.d. assumption.
- The Information Vacuum: In a "Cold Start" scenario, there is literally zero data in the target network to build a prediction.
Methodology: The Power of Anchor Links
The core insight of this paper is the Anchor Link. Many users link their social accounts (e.g., "Sign in with Twitter"). These links act as bridges across heterogeneous networks.
1. Within-Network Personalized Sampling (PS)
To solve the distribution gap, the authors don't just use all old users. They use a Sampling Vector () to extract a sub-network of old users that "looks like" the new users. The objective function maximizes:
- Relevance: How similar are these old users to the new user cohort?
- Diversity: Ensuring the sampled data isn't one-dimensional.
- Structure Maintenance: Preserving the essential topology of the social graph.
2. Cross-Network Feature Fusion
Once aligned, SCAN-PS extracts features from both networks. For a potential link , it looks at:
- Intra-network features: Behavior on the target platform.
- Inter-network features: Behavior on the source platform (the "Pseudo Label").
Fig 1: Using Anchor Links to transfer knowledge between Twitter and Foursquare.
Experimental Battleground
The researchers tested this on real-world datasets from Foursquare and Twitter.
Performance in Cold Start
When the "Remaining Information Ratio" is 0.0 (meaning the user has 0 posts and 0 friends on the target platform), traditional methods like Jaccard Coefficient are no better than a random guess (AUC 0.5). SCAN-PS, however, hits an AUC of 0.783 because it effectively "teleports" the user's social preferences from the source network.
Table 1: SCAN-PS consistently outperforms baselines across different levels of "newness".
Critical Insight & Conclusion
The genius of SCAN-PS lies in its admission that local data is not enough. By using Transfer Learning via Anchor Links, the "Cold Start" problem is transformed from a missing data problem into a data alignment problem.
Key Takeaways for Practitioners:
- Sampling Matters: Don't train your recommendation engines on "Power Users" and expect them to work for newcomers. Use Personalized Sampling to bridge the gap.
- The Multi-Network Identity: Encouraging users to link other social accounts is the most effective way to solve the cold-start problem in modern apps.
Limitations: The method relies on the existence of Anchor Links. For users who value anonymity and don't cross-link accounts, the model's advantage diminishes. Future work likely involves "Implicit Alignment," where users are linked based on behavior patterns rather than explicit IDs.
