Identifying the Stranger: Privacy-Preserving Professional Profiling via Mobile Big Data
Identifying unfamiliar callers’ professions from privacy-preserving mobile phone data
This paper introduces a machine learning framework to identify the professions of unfamiliar callers (Normal, Taxi Driver, Delivery, or Telemarketer/Fraudster) using privacy-preserving mobile cellular data. By analyzing anonymized web service requests from 1,282 users in Shanghai, the authors achieved an overall classification accuracy of 75.64% using a Random Forest model.
TL;DR
In an era of rampant telemarketing and "last-mile" delivery services, knowing who is calling is a matter of safety and efficiency. This paper proposes a system that identifies a caller's profession—Taxi Driver, Delivery Staff, or Telemarketer—by analyzing statistical patterns in their mobile data traffic. By focusing on "how" people move and use apps rather than "where" they are or "what" they specifically do, the method achieves over 75% accuracy while remaining strictly GDPR-compliant.
The "Lazy Label" Problem: Why Current Systems Fail
Most caller ID apps (like TrueCaller or Baidu's label service) depend on users manually flagging numbers. This creates significant issues:
- Inaccuracy: One disgruntled user can mislabel a normal person as "Harassment."
- Latency: A fraudster can make thousands of calls before enough reports accumulate to trigger a public label.
- Data Sensitivity: Previous research attempts to automate this often required "raw" data—exact GPS coordinates or full App usage logs—which are now restricted by privacy laws like GDPR.
Methodology: The Art of Statistical Fingerprinting
The authors' core insight is that professions have distinct behavioral signatures that persist even after sensitive details are stripped away. They propose three primary feature sets:
1. Directional Mobility Patterns
Instead of tracking exact locations, the model calculates the Standard Deviation (SD) of a user's position across 12 evenly distributed directions.
- Taxi Drivers: High SD in all directions (wide, multi-directional coverage).
- Delivery Staff: High SD in specific regions (hub-and-spoke patterns).
- Telemarketers: Simple strip-like patterns (basic home-to-office commuting).
Figure 1: Comparison of location distributions (a) and 12-direction SD vectors (b) for different professions.
2. Request Volume & Temporal Activeness
By dividing the day into six slices (6:00 to 24:00), the model tracks the volume of web service requests. Telemarketers and drivers show higher activity after 18:00, while delivery staff exhibit a more stable, evenly distributed request pattern throughout the day.
3. App Preference Distribution
To maintain privacy, the model doesn't look at which app is used, but rather the distribution curve of the top 10 most-used domains. Normal users have a "flat" distribution (diverse interests), whereas professional users have a "steep" curve dominated by 1 or 2 work-specific apps (e.g., driver versions of Uber/DiDi).
Figure 2: Sorted usage rates demonstrating the concentrated app focus of professionals vs. normal users.
Experiments and Results
The study utilized a real-world dataset from a major Chinese telecom operator in Shanghai, covering 1,282 users.
- Top Performer: Random Forest (RF) achieved the highest overall accuracy of 75.64%.
- Professions: Identification was most accurate for Drivers (79.12%) and Delivery staff (78.84%).
- Efficiency: The researchers found that data from just one day was sufficient to reach these accuracy levels, making real-time identification feasible for newly activated professional numbers.
Table 1: Accuracy results across different machine learning algorithms.
Critical Insight: Why This Matters
The breakthrough here isn't just the 75% accuracy—it's the inductive bias of the feature engineering. By using sorted directional SDs, the model becomes rotation-invariant. Whether a delivery driver works in a north-south grid or a circular city layout, their "distribution" remains identifiable.
Limitations:
- The model struggles slightly to distinguish "Normal" users from "Harassment" callers (both show lower mobility compared to drivers).
- Professional shifts (people changing jobs) can introduce noise into the labels over a 3-month period.
Conclusion
This work bridges the gap between big data utility and user privacy. It proves that we don't need to know exactly where a person is to understand their professional role; their movement and digital rhythms speak for themselves. This paves the way for carrier-level identification services that protect users from fraud without compromising the privacy of the callers.
