TSU: Leveraging Social Trust for High-Precision Home Location Identification

We Know Where You Are: Home Location Identification in Location-Based Social Networks

2016-08-01
Yulong Gu, Yuan Yao, Weidong Liu, Jiaxing Song
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces TSU (Trust-based Unified Model), a probabilistic framework for Home Location Identification in Location-Based Social Networks (LBSNs). By integrating social friendship, check-in data, and a novel "social trust" metric, the method identifies user home locations even with sparse data, achieving a SOTA accuracy of up to 92.1% on users with check-ins.

TL;DR

Researchers from Tsinghua University have developed TSU (Trust-based Unified Model), a probabilistic framework designed to pin down user home locations in Location-Based Social Networks (LBSNs) like Foursquare. By introducing "social trust"—a measure of social structural closeness—and a two-stage iterative algorithm, the model achieves a 92.1% accuracy for active users and significantly outperforms existing baselines (by up to 47.4% in sparse data environments).

Background: The Sparsity Challenge

Knowing a user's home location is the "Holy Grail" for personalized services, localized news, and targeted advertising. However, the data is notoriously messy:

  • Privacy: Most users leave their profile location blank.
  • Noise: Users check in at tourist spots far from home.
  • Sparsity: A vast majority of users have zero or very few check-ins.

While previous works used either Content-Based (analyzing tweets) or Check-in Based approaches, they often treated every friend as equal. In reality, a "friend" you share 50 mutual connections with (high social trust) is a much better indicator of your location than a random celebrity you follow.

Methodology: The TSU Model

The core innovation of TSU is the integration of Social Trust into an Influence Model.

1. Defining Social Trust

Instead of a binary 0/1 friendship, TSU calculates the Jaccard Similarity between user friend sets: The intuition? People with more common friends are structurally "closer" and, statistically, geographically closer.

2. The Influence Distribution

TSU models every user and venue as having an "influence scope" represented by a bivariate Gaussian distribution. The probability of an edge (friendship or check-in) is calculated based on the distance between the tail node (the person) and the center of the head node's influence (the friend's home or the venue).

3. HLIA: Two-Stage Algorithm

The identification process, named HLIA (Home Location Identification Algorithm), works in two phases:

  • Stage 1 (Initialization): For users with check-ins, a Single-pass Clustering (SPClustering) identifies the largest cluster of activity to set an initial home.
  • Stage 2 (Iterative Optimization): A global iteration updates the influence scope () and coordinates () for users without data, maximizing the joint likelihood of the entire social-spatial graph.

Overall Strategy Fig 1: Heterogeneous Graph Representation of LBSN interactions.

Experiments & SOTA Performance

The authors tested TSU on a massive Foursquare dataset involving 835,896 users and 12.9 million social edges in the US.

Key Findings:

  • Efficiency in Sparsity: Even with an average of only 2.7 check-ins per user, TSU reached 92.1% accuracy.
  • Iterative Superiority: Against the previous SOTA model (UDI), TSU showed a 6.9% improvement overall.
  • Robustness: As the pool of "known" users shrinks (down to 20%), TSU’s performance lead grows to a massive 47.4% improvement over UDI, proving its ability to propagate location info through the social trust graph effectively.

Performance Comparison Fig 2: Accuracy vs. Error Distance. TSU maintains higher accuracy across all distance thresholds compared to UDI.

Critical Insight: Why Does It Work?

The effectiveness of TSU stems from two departures from prior work:

  1. Weighted Influence: By using social trust, the model ignores "weak ties" that lead to geographical outliers and focuses on "strong ties" likely to be within the same city.
  2. Stability: By not re-updating check-in-based initializations in the final iteration (unless necessary), the model avoids "washing out" high-confidence check-in data with noisier social predictions.

Conclusion & Future Work

TSU demonstrates that the structure of our social circles is a powerful "coordinate system" in its own right. While the current model is highly effective, the authors suggest that adding temporal information (e.g., distinguishing between a daytime workplace check-in and a nighttime home check-in) could further push the boundaries of LBSN profiling.

Takeaway for Practitioners: When building location-aware recommenders, look beyond the raw check-in; the "trust" inferred from mutual social connections is often the missing piece of the puzzle.


Limitations

  • Computation: Iterative global optimization on graphs with millions of nodes is computationally expensive.
  • Dynamic Privacy: As privacy settings on LBSNs evolve, the availability of friend lists (the basis for Jaccard Similarity) may decrease, potentially weakening the "social trust" signal.

Find Similar Papers

Try Our Examples

  • Search for recent papers that integrate temporal patterns or "check-in time of day" into probabilistic home location identification models.
  • Which studies first introduced the use of Jaccard Similarity on social graphs to refine geographic localization, and how does TSU's implementation differ?
  • Explore how trust-based influence models are being adapted for privacy-preserving location recommendation without exposing specific user coordinates.
Contents
TSU: Leveraging Social Trust for High-Precision Home Location Identification
1. TL;DR
2. Background: The Sparsity Challenge
3. Methodology: The TSU Model
3.1. 1. Defining Social Trust
3.2. 2. The Influence Distribution
3.3. 3. HLIA: Two-Stage Algorithm
4. Experiments & SOTA Performance
4.1. Key Findings:
5. Critical Insight: Why Does It Work?
6. Conclusion & Future Work
6.1. Limitations