Protecting the Individual: Privacy in Topic-Aware Social Influence Networks

Individual privacy in social influence networks

2015-12-28
Sara Hajian, Tamir Tassa, Francesco Bonchi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel framework for preserving individual privacy in topic-dependent social influence networks. It proposes a randomization method combining edge sparsification and weight reduction to achieve "k-obfuscation," successfully maintaining high utility for influence maximization tasks on real-world datasets like Flixster and Digg.

TL;DR

In the age of viral marketing, social networks are no longer just "who knows whom"—they are "who influences whom, on what topic, and by how much." This paper tackles the immense privacy risk of publishing such rich data. By applying strategic randomization (edge deletion and weight scaling), the authors achieve k-obfuscation, ensuring users "blend into the crowd" while keeping the data 80-90% effective for identifying top influencers.

The Problem: The Curse of Rich Data

Most privacy research treats social networks as simple dots and lines. In reality, platforms like Twitter or Flixster track topic-dependent influence. If an adversary knows you follow 10 people and are specifically a "70% influence" on music but "10% influence" on politics, that unique signature acts as a fingerprint.

Prior work in k-anonymity fails here because of the Curse of Dimensionality. Trying to make everyone’s "influence profile" look exactly like other people requires changing so much data that the resulting graph becomes useless for researchers.

Methodology: Strategic Randomization

The authors move away from deterministic "bucketizing" and embrace controlled randomness through two main mechanisms:

  1. Edge Sparsification: For every edge, a coin is flipped (Bernoulli trial). With probability , the edge is deleted. This breaks the structural "fingerprint" of the node's degree.
  2. Weight Reduction: Topic weights are multiplied by a random factor . To preserve utility, the authors use a linearly increasing distribution , making it more likely to choose a factor close to 1 (minimal change) than 0 (total erasure).

Proposed Method Process

Formalizing Privacy: k-Obfuscation

Rather than saying "every node must look identical," the authors use Entropy. If an adversary looks at the perturbed graph and tries to find "Alice," the entropy of their probability distribution must be at least . Essentially, the adversary should be as confused as if they were looking at equally likely candidates.

Experiments: Does it stay useful?

The authors tested their method on Flixster (movie ratings) and Digg (social news). They measured the utility via Topic-aware Influence Maximization (TIM)—the task of finding the best 50 users to start a viral campaign.

Key Performance Metrics:

  • Structural Integrity: Graph properties like diameter and clustering coefficients remained stable under moderate (edge removal).
  • Query Precision: On the Digg dataset, even with significant perturbation, the system identified 8 out of the top 10 most influential users correctly.
  • Obfuscation Gain: While the original graph offered almost no privacy (most nodes were unique), the perturbation successfully hid the majority of users within a "crowd" of or more.

Influence Distribution

Critical Analysis

The beauty of this approach is its simplicity and efficiency. While Differential Privacy (DP) is often considered the gold standard, implementing DP for graph structures often results in massive noise that destroys edge-level utility. This randomization approach provides a practical middle ground for data publishers.

Limitations: The "1-neighborhood" adversarial assumption is strong, but what if the adversary knows 2-hop or 3-hop patterns? The authors acknowledge that more complex structural patterns could still pose risks, and future work is needed to extend these protections to more global structural knowledge.

Conclusion

This work proves that we don't have to choose between rich social insights and individual privacy. By carefully "blurring" the edges and weights of an influence network, we can protect the identities of users while still allowing advertisers and sociologists to understand the flow of information across topics.

Seed Set Precision

Find Similar Papers

Try Our Examples

  • Find recent papers addressing identity re-identification attacks in multi-layer or heterogeneous social networks beyond simple graph structures.
  • Which study first introduced the concept of "k-obfuscation" using entropy as a privacy metric, and how does this paper's application to topic models extend that theory?
  • What are the latest differential privacy mechanisms designed for directed graphs with continuous edge weights in the context of influence propagation?
Contents
Protecting the Individual: Privacy in Topic-Aware Social Influence Networks
1. TL;DR
2. The Problem: The Curse of Rich Data
3. Methodology: Strategic Randomization
3.1. Formalizing Privacy: k-Obfuscation
4. Experiments: Does it stay useful?
4.1. Key Performance Metrics:
5. Critical Analysis
6. Conclusion