Beyond the Centralized Silo: Ontology-Enabled Privacy in the di.me Ecosystem

Ontology-Enabled Access Control and Privacy Recommendations

2014-12-24
Marcel Heupel, Lars Fischer, Mohamed Bourimi, Simon Scerri
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents the di.me project, a decentralized social networking system that utilizes an ontology-based framework to provide user-controlled access control and privacy recommendations. By leveraging semantic technologies and NLP, it detects identity linkability and warns users about unintended information disclosure in live streams.

TL;DR

The di.me project introduces a decentralized "userware" that replaces centralized social networks with a semantic-driven personal server. By integrating multiple online identities into a single Personal Information Model (PIM), it uses Natural Language Processing (NLP) and weighted semantic matching to warn users about identity linkability risks and accidental data leaks in their status updates.

The Linkability Crisis in Modern OSNs

In our current digital landscape, we maintain fragmented identities across LinkedIn, Facebook, and Twitter. While we might attempt to keep a "private" persona separate from a "professional" one, research shows that re-identification rates remain alarmingly high—often exceeding 88%.

The core problem is the lack of semantic awareness. Most platforms treat data as isolated strings; they don't understand that a pseudonym on Flickr and a real name on Twitter both point to the same physical entity. Furthermore, users often post "LivePosts" (microblogs) that contain sensitive context—like tagging a friend who is supposed to be at work—without realizing the privacy implications for others.

Methodology: The Semantic Core

The di.me architecture is built on a two-layer control mechanism that decouples the semantic core from the hosting environment. This ensures that even if you migrate your data to a new server, your privacy rules stay with the data.

1. Semantic Equivalence Detection

To solve the linkability issue, di.me employs a four-step matching pipeline:

  • Linguistic Analysis: Breaking down complex strings (e.g., extracting "Bonn" from a full address).
  • Syntactic Matching: Using Monge and Elkan recursive algorithms to find string similarities.
  • Semantic Expansion: Linking extracted entities (like a company name) to Global Knowledge Bases like DBPedia or the user's private PIM.
  • Weighted Scoring: Assigning higher importance to "Inverse Functional Properties" like email addresses over generic attributes like "Country."

Overall architecture of the di.me system

2. Privacy-Aware NLP for Live Streams

The system doesn't just look at profile settings; it "reads" your posts before they go live. Using the Live Post Ontology (DLPO), the system decomposes a post into its constituent parts: Image, Check-In, and Status.

If the NLP engine detects that you are mentioning a contact ("Anna") alongside a location ("Beach") that contradicts Anna's known private schedule, the system triggers a Privacy Recommendation.

Information Extraction Pipeline for Privacy Advisory

Experimental Validation

The system was validated through large-scale user trials. Key findings include:

  • Precision: The semantic matching reached 82% accuracy in identifying equivalent profiles across different networks.
  • User Acceptance: 81% of users found the privacy recommendations valuable, highlighting a strong market demand for "Privacy Advisors" in the CRM and Social sectors.

The comparison with other Information Extraction (IE) techniques shows that di.me is one of the few systems that targets all four major named entities (People, Events, Activities, Locations) and links them back to a personal knowledge base for context-aware reasoning.

Comparison of IE techniques on microposts

Critical Insight & The "Sticky Policy" Future

The most impactful contribution of this work is the concept of "Sticky Policies." By attaching ontology-based metadata to shared items, the original owner's intent (e.g., "do not re-share") follows the file even after it enters another person's userware.

While the current implementation relies on user cooperation and decentralized protocols, the logic is highly portable. Even centralized giants could adopt these "Semantic Privacy Guards" to provide more transparent control to their users.

Conclusion

The di.me project moves us away from passive privacy (simply checking boxes) toward proactive privacy. By understanding the meaning of our data through ontologies and NLP, the system becomes an intelligent agent acting in the user's best interest, preventing the "unintentional linkability" that current social media platforms exploit.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Knowledge Graphs and Ontology-based Access Control (OBAC) to prevent cross-platform identity linkage in decentralized web environments.
  • Which paper originally proposed the 'Privacy Preference Ontology (PPO)', and how does the di.me project extend its application to 'closed-world' decentralized systems?
  • What are the current state-of-the-art Natural Language Processing (NLP) techniques for detecting sensitive information (PII) leakage in real-time social media streams beyond named entity recognition?
Contents
Beyond the Centralized Silo: Ontology-Enabled Privacy in the di.me Ecosystem
1. TL;DR
2. The Linkability Crisis in Modern OSNs
3. Methodology: The Semantic Core
3.1. 1. Semantic Equivalence Detection
3.2. 2. Privacy-Aware NLP for Live Streams
4. Experimental Validation
5. Critical Insight & The "Sticky Policy" Future
6. Conclusion