Beyond the Lone Traveler: Deanonymizing Mobility Traces via Co-Location Information

Deanonymizing mobility traces with co-location information

2017-10-01
Youssef Khazbak, Guohong Cao
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the identity inference problem in anonymized mobility traces by exploiting co-location side information from social networks like Twitter and Swarm. The authors develop two maximum likelihood-based attack frameworks that utilize location and co-location data to deanonymize users with high accuracy, even when traces are obfuscated.

TL;DR

Think deleting your name and adding noise to your GPS coordinates keeps you anonymous? Think again. This paper reveals that co-location information—the "I'm here with Blanca" check-ins on Swarm/Twitter—creates a web of dependencies that allows attackers to deanonymize "hidden" traces. By modeling these social links via MLE and Hidden Markov Models, an adversary can identify your full movement history with up to 80% accuracy, even if you never personally post your location.

The "Interdependent Privacy" Gap

Most privacy research treats users as islands. We assume that if we obfuscate User A's data, User A is safe. However, in the age of Online Social Networks (OSNs), privacy is interdependent.

The authors identify a critical vulnerability: even if you are highly privacy-conscious, a friend checking in with you at a restaurant leaks your location. More importantly, this single point of shared truth can be used as an "anchor" to link your entire anonymized mobility trace (a sequence of timestamps and coordinates) to your real identity.

Methodology: The Math of Social Tracking

The paper formalizes the threat into two distinct attack vectors based on the adversary's knowledge.

1. The MLE Attack (Location + Co-location)

When an attacker has access to some of your reported locations () and co-location events (), they use a Maximum Likelihood Estimator.

  • The Intuition: The attacker looks at all available anonymized traces and asks: "Which trace maximizes the probability of seeing these specific check-ins and co-locations?"
  • Handling Noise: If the data is obfuscated with Gaussian noise, the attack transforms into a least-squares optimization problem, essentially finding the "path of best fit" across multiple traces simultaneously.

2. The HMM Attack (Mobility Profiles)

If the attacker also knows your past habits (e.g., you usually go from home to a specific gym), they build a Markov Transition Matrix.

  • The Intuition: The problem becomes a Hidden Markov Model (HMM). The hidden states are your true locations, while the anonymized traces and social check-ins are the observations.
  • The Forward Algorithm: By using an iterative forward algorithm, the attacker can calculate the likelihood of a trace even if the observed data is extremely sparse ( direct location samples).

Model Architecture and Motivation Fig 1: Motivation - Adding co-location information (Alice with Bob at t3) allows an adversary to distinguish between traces that otherwise look identical.

Experimental Proof: Taxis and Buses

The authors tested their theories on two high-entropy datasets: Roma Taxis and Shanghai Buses.

Key Findings:

  • The Power of 'With': Identication accuracy scales aggressively with co-location count. In the taxi dataset, accuracy jumps significantly once co-locations are introduced as side information.
  • Regularity as a Weakness: The attack performed even better on the bus dataset because bus routes are highly regular, making the Markovian mobility profiles extremely accurate.
  • Trace Volume vs. Privacy: As the number of candidate traces () increases, the "crowd" offers some protection, but co-location information acts as a powerful filter that quickly narrows down the possibilities.

Accuracy vs Side Information Fig 2: Identification accuracy increases as the number of observed locations (k) and co-locations (c) grows, showing that social data provides a massive boost to deanonymization.

Critical Analysis & Conclusion

Takeaway

The core contribution of this work is the rigorous proof that social context is as dangerous as spatial data. By leveraging the "I'm here with..." feature of modern apps, adversaries can bypass sophisticated spatial cloaking and k-anonymity defenses.

Limitations

  • Computational Complexity: The HMM approach is . While feasible for hundreds of traces, it may require optimization for city-scale datasets with millions of users.
  • Data Availability: The attack assumes the adversary can link social media handles to the entities in the mobility traces, which is a significant (though often realistic) hurdle.

Future Outlook

This paper serves as a warning for the designers of Location-Based Services (LBS). Future privacy-preserving algorithms must be socially-aware, perhaps by implementing "Group-based Differential Privacy" where the privacy budget is calculated based on the entire social cluster rather than just the individual.

Find Similar Papers

Try Our Examples

  • Find recent papers that propose social-aware location privacy-preserving mechanisms (LPPM) specifically designed to mitigate co-location information leakage.
  • Which study first introduced the concept of 'interdependent privacy' in location-based services, and how does this paper quantify that risk compared to the original work?
  • Explore how graph neural networks (GNNs) have been applied to the problem of deanonymizing mobility traces using social relationship graphs as side information.
Contents
Beyond the Lone Traveler: Deanonymizing Mobility Traces via Co-Location Information
1. TL;DR
2. The "Interdependent Privacy" Gap
3. Methodology: The Math of Social Tracking
3.1. 1. The MLE Attack (Location + Co-location)
3.2. 2. The HMM Attack (Mobility Profiles)
4. Experimental Proof: Taxis and Buses
4.1. Key Findings:
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook