What You Want Is Not What You Get: Using Machine Learning to Fix Facebook Privacy
13078_What you want is not what you get predicting sharing policies for text-based content on facebook.
The paper introduces an automated approach using machine learning, specifically the MaxEnt classifier, to predict privacy sharing policies for text-based Facebook posts. By analyzing post content and metadata, the method achieves an average accuracy of 81%, significantly outperforming Facebook's default "previous policy" baseline.
TL;DR
Social media users are notoriously bad at managing privacy settings, often sharing sensitive text with the wrong audience. This research moves beyond Facebook's primitive "repeat the last setting" default by using a MaxEnt (Maximum Entropy) classifier to predict intended sharing policies. It achieves 81% accuracy on average, and up to 94% for consistent users, effectively halving the rate of privacy misconfigurations.
The "Default" Trap: Why Privacy Fails
Most users operate under a "momentary inattentiveness" during social media interactions. Facebook’s native solution—defaulting a new post’s privacy to whatever you used last—is a double-edged sword. While convenient, it ignores the content of the post. An update about a Friday night party requires a different audience than a professional link shared on a Monday morning.
The authors' survey revealed a startling "Confidence Factor": many users' implemented policies (what they actually clicked) did not match their intended policies (what they actually wanted). This "noise" makes automated assistance both difficult and necessary.
Methodology: Beyond Simple Keyword Matching
The core of the solution is a MaxEnt classifier. Unlike simpler models, MaxEnt doesn't assume anything about the data it hasn't seen, making it robust for the diverse and "messy" nature of social media text.
Key Features for Prediction:
- Textual Content: Filtering stop words and transforming emoticons (e.g., ":)" becomes "happy") to capture sentiment.
- Temporal Bucketing: Categorizing posts into "Office hours," "Evenings/Weekends," and "Nights."
- Metadata: Binary flags for attachments (images/videos) and domain-specific URL parsing.
- Structural Context: Using word bi-grams (e.g., "Sherlock Holmes" vs "Sherlock") to capture deeper meaning.
(Note: Refer to Table 5 in the paper for detailed feature exclusion impacts)
Critical Results: Smarter than the Average Default
The researchers tested their model against several datasets, most notably the "Pruned Clean" dataset, which focused on data where the user's true intent was verified.
| Strategy | Avg. Intended Accuracy |
|---|---|
| Facebook Default (Previous Policy) | 67% |
| MaxEnt (Our Approach) | 81% |
| Top-Two Suggestion (Exact or 2nd choice) | 94% |
The reduction in error is significant: a 45% decrease in misconfigured posts. Perhaps most interestingly, the model requires very little "warm-up" time. As seen in the paper's learning curve, the accuracy stabilizes after just a few initial training samples.
Figure: Accuracy ramps up quickly as training data grows.
Deep Insight: The Feedback Loop
The most profound takeaway is the correlation between user consistency and machine accuracy. For users who generally know what they want (high confidence), the machine acts as a perfect shield (94%+ accuracy). For "noisy" users, the machine struggles.
However, the authors argue for a virtuous cycle: if a social network implements a tool like this, users will set better policies initially. This "cleaner" data then feeds back into the algorithm, making the predictions even more accurate over time.
Limitations and The Future
While the study focused on text, modern social media is visual. The authors acknowledge that integrating image analysis (CV) with their text-based MaxEnt model is the logical next step. Furthermore, the sample size (42 participants) is small by industry standards, but the strength of the statistical correlation suggests the findings would hold at scale.
Conclusion
We no longer have to rely on "last-used" defaults that expose our private lives to the public. By treating privacy settings as a classification problem, we can build "Privacy Assistants" that understand the nuance of what we say and when we say it.
