TSU: Leveraging Social Trust for High-Precision Home Location Identification
We Know Where You Are: Home Location Identification in Location-Based Social Networks
The paper introduces TSU (Trust-based Unified Model), a probabilistic framework for Home Location Identification in Location-Based Social Networks (LBSNs). By integrating social friendship, check-in data, and a novel "social trust" metric, the method identifies user home locations even with sparse data, achieving a SOTA accuracy of up to 92.1% on users with check-ins.
TL;DR
Researchers from Tsinghua University have developed TSU (Trust-based Unified Model), a probabilistic framework designed to pin down user home locations in Location-Based Social Networks (LBSNs) like Foursquare. By introducing "social trust"—a measure of social structural closeness—and a two-stage iterative algorithm, the model achieves a 92.1% accuracy for active users and significantly outperforms existing baselines (by up to 47.4% in sparse data environments).
Background: The Sparsity Challenge
Knowing a user's home location is the "Holy Grail" for personalized services, localized news, and targeted advertising. However, the data is notoriously messy:
- Privacy: Most users leave their profile location blank.
- Noise: Users check in at tourist spots far from home.
- Sparsity: A vast majority of users have zero or very few check-ins.
While previous works used either Content-Based (analyzing tweets) or Check-in Based approaches, they often treated every friend as equal. In reality, a "friend" you share 50 mutual connections with (high social trust) is a much better indicator of your location than a random celebrity you follow.
Methodology: The TSU Model
The core innovation of TSU is the integration of Social Trust into an Influence Model.
1. Defining Social Trust
Instead of a binary 0/1 friendship, TSU calculates the Jaccard Similarity between user friend sets: The intuition? People with more common friends are structurally "closer" and, statistically, geographically closer.
2. The Influence Distribution
TSU models every user and venue as having an "influence scope" represented by a bivariate Gaussian distribution. The probability of an edge (friendship or check-in) is calculated based on the distance between the tail node (the person) and the center of the head node's influence (the friend's home or the venue).
3. HLIA: Two-Stage Algorithm
The identification process, named HLIA (Home Location Identification Algorithm), works in two phases:
- Stage 1 (Initialization): For users with check-ins, a Single-pass Clustering (SPClustering) identifies the largest cluster of activity to set an initial home.
- Stage 2 (Iterative Optimization): A global iteration updates the influence scope () and coordinates () for users without data, maximizing the joint likelihood of the entire social-spatial graph.
Fig 1: Heterogeneous Graph Representation of LBSN interactions.
Experiments & SOTA Performance
The authors tested TSU on a massive Foursquare dataset involving 835,896 users and 12.9 million social edges in the US.
Key Findings:
- Efficiency in Sparsity: Even with an average of only 2.7 check-ins per user, TSU reached 92.1% accuracy.
- Iterative Superiority: Against the previous SOTA model (UDI), TSU showed a 6.9% improvement overall.
- Robustness: As the pool of "known" users shrinks (down to 20%), TSU’s performance lead grows to a massive 47.4% improvement over UDI, proving its ability to propagate location info through the social trust graph effectively.
Fig 2: Accuracy vs. Error Distance. TSU maintains higher accuracy across all distance thresholds compared to UDI.
Critical Insight: Why Does It Work?
The effectiveness of TSU stems from two departures from prior work:
- Weighted Influence: By using social trust, the model ignores "weak ties" that lead to geographical outliers and focuses on "strong ties" likely to be within the same city.
- Stability: By not re-updating check-in-based initializations in the final iteration (unless necessary), the model avoids "washing out" high-confidence check-in data with noisier social predictions.
Conclusion & Future Work
TSU demonstrates that the structure of our social circles is a powerful "coordinate system" in its own right. While the current model is highly effective, the authors suggest that adding temporal information (e.g., distinguishing between a daytime workplace check-in and a nighttime home check-in) could further push the boundaries of LBSN profiling.
Takeaway for Practitioners: When building location-aware recommenders, look beyond the raw check-in; the "trust" inferred from mutual social connections is often the missing piece of the puzzle.
Limitations
- Computation: Iterative global optimization on graphs with millions of nodes is computationally expensive.
- Dynamic Privacy: As privacy settings on LBSNs evolve, the availability of friend lists (the basis for Jaccard Similarity) may decrease, potentially weakening the "social trust" signal.
