The Path to Standardized Privacy: Beyond the "I Agree" Wall
Investigating Similarity Between Privacy Policies of Social Networking Sites as a Precursor for Standardization
This paper investigates the feasibility of standardizing privacy policies for Social Networking Sites (SNS) through a comparative analysis of policies from Facebook, Twitter, LinkedIn, Google+, Pinterest, and Flickr. Using thematic analysis and Cross-Document Structure Theory (CST), the authors evaluate similarity based on specific clauses and compliance with 40 UK ICO recommendations, proposing a roadmap for industry-wide standardization.
TL;DR
Privacy policies are notoriously "tl;dr" for users, but are they also fundamentally different under the hood? This research breaks down the internal architecture of privacy policies from the world's top 6 social networking sites (SNS). It reveals a striking paradox: while the legal wording varies wildly, the actual information content is remarkably similar. This finding argues that standardization is not only possible but overdue.
Problem & Motivation: The Illusion of Choice
The "Notice and Choice" model of online privacy is broken. We are forced to navigate a "Risk Society" where our ethnic origin, religious beliefs, and sexual orientation can be predicted just from Facebook "likes." Despite this, privacy policies remain:
- Inaccessible: Too long and laden with legalese.
- Inconsistent: A "delete" button on one site might mean "permanent erasure," while "close account" on another might mean "deactivation."
- Disincentivized: SNS benefit from data harvesting, and the legal complexity of multi-jurisdictional compliance makes "concise yet compliant" policies a unicorn.
The authors argue that standardization is the prerequisite for all other improvements, including visualization tools and machine-readable privacy assistants.
Methodology: Deconstructing the Legalese
To test if standardization is a pipe dream, the researchers scrutinized the "first layer" policies of Facebook, Twitter, LinkedIn, Google+, Pinterest, and Flickr.
The Two-Pronged Attack
- Clause-Level Similarity: Granular comparison of the actual sentences/clauses used.
- Thematic-Level Similarity: Do the policies cover the 40 recommendations set by the UK Information Commissioner’s Office (ICO)?
They utilized Cross-Document Structure Theory (CST) to identify relationships between segments and calculated the Jaccard Similarity Coefficient to quantify just how much these legal documents overlapped.
Table: Comparison of total clauses and removal of redundant information across SNS.
Experiments & Results: The "Clause vs. Theme" Paradox
The results were polarized.
1. Low Literal Similarity
At the sentence level, the policies were quite different. The average similarity was a mere 15%. The reasons?
- Functionality: Facebook's "Instant Personalization" vs. LinkedIn's "Polls."
- Semantics: Interchanging "close" and "deactivate."
- Elaboration: Some sites provide physical addresses and detailed "cookie" definitions; others don't.
2. High Thematic Alignment
When looking at what they actually told the users (e.g., "how long we keep data"), similarity jumped to 52% - 81%. As shown in the graph below, most recommendations were either addressed by all six SNS or by none of them—showing high collective behavior.
Graph: Distribution of ICO recommendation coverage, showing a trend where themes are either universally adopted or ignored.
Deep Insight: A Roadmap to Standardization
The study concludes that the "DNA" of SNS privacy policies is already similar enough to support a standard. The authors offer five strategic recommendations:
- Theme-First Approach: Start by standardizing the topics (headers) rather than the sentences.
- Modular Functionality: Separate "General functionality" (logs, cookies) from site-specific features.
- Unified Definitions: Standardize technical glossaries so a "cookie" means the same thing everywhere.
- Semantic Synchronization: Force a standard meaning for "delete," "close," and "deactivate."
- Evidence of Absence: Make it clear when a theme is not addressed and why (transparency of omission).
Conclusion & Future Outlook
This work shifts the conversation from "privacy policies are too hard to read" to "privacy policies are structurally homogeneous enough to be automated." By aligning themes first, we pave the way for a future where a browser plugin can instantly compare the "Privacy Nutrition Label" of any two sites. While current literal similarity is low, the thematic core is stable—meaning the legal "bloat" is the only thing standing between users and a transparent Web 2.0.
