SSDMV: Shielding Social Networks with Deep Multi-View Fusion and Semi-Supervised Learning

SSDMV: Semi-Supervised Deep Social Spammer Detection by Multi-view Data Fusion

2018-11-01
Chaozhuo Li, Senzhang Wang, Lifang He, Philip S. Yu, Yanbo Liang, Zhoujun Li
Summary
Problem
Method
Results
Takeaways
Abstract

SSDMV is a novel semi-supervised deep learning framework for social spammer detection that fuses multi-view data (relations, demographics, text, and symbols). It achieves state-of-the-art performance by combining a Correlated-Ladder-Network (CLN) for robust feature learning with a label inference module, effectively handling annotation scarcity.

TL;DR

Social media platforms are under constant siege by spammers. SSDMV (Semi-Supervised Deep Multi-View) is a breakthrough framework that solves the "label scarcity" problem. By leveraging a Correlated-Ladder-Network, it learns highly discriminative user features from relations, text, and metadata simultaneously. Remarkably, it requires 98% less labeled data than previous supervised SOTA methods to achieve the same detection accuracy.

Problem & Motivation: The "Shallow" Gap and Label Scarcity

Detecting spammers is a cat-and-mouse game. While supervised learning works well, annotating millions of users is economically impossible. Conversely, unsupervised methods suffer from high false-positives.

The authors identified three critical gaps in existing research:

  1. Non-Linearity: Social data (language and networks) is highly complex; shallow models can't "see" the nuanced patterns of bot behavior.
  2. Loose Coupling: Most models just concatenate features from different views (e.g., Text + Relations), ignoring that these views are deeply correlated.
  3. Task Irrelevance: Feature learning and classification are often separated, meaning the model might learn features that don't actually help distinguish spammers from humans.

Methodology: The Correlated-Ladder-Network (CLN)

The heart of SSDMV is the Correlated-Ladder-Network. Unlike a standard auto-encoder, a Ladder Network uses a denoising mechanism and skip connections to maintain structural integrity while learning abstract features.

The Innovation: The Filter Gate

The authors introduced a Filter Gate to allow information to "flow" between views. If a user's tweet history is sparse, the model can "borrow" certainty from their social relation view. This gate acts as an intelligent switch, determining how much cross-view information is beneficial versus how much is just noise.

Overall Architecture Fig 1: The SSDMV framework, showing the interaction between the CLN feature learner and the MLP label inference module.

Unified Optimization

SSDMV optimizes two things at once:

  • Reconstruction Loss: Keeps the features representative of the original data (using both labeled and unlabeled users).
  • Classification Loss: Ensures the features are actually useful for spotting spammers (using only labeled users).

Experiments: Doing More with Less

The results are staggering. In a head-to-head comparison on a Twitter dataset, SSDMV reached a 0.915 F1-Score using only 200 labeled samples. For comparison, the supervised MVSD method needed 16,000 labels to reach a similar performance (0.911).

Performance Comparison Table 1: Performance comparison across different label ratios. SSDMV consistently leads even when labels are extremely scarce.

Robustness to Noise and Sparsity

The model was tested by intentionally deleting 80% of a user's social links or tweets. Because of the Multi-View Fusion, SSDMV was able to compensate for the missing data in one view by utilizing the others, effectively "filling in the blanks" of malicious behavior.

Deep Insights: Visualizing the Latent Space

The superiority of the model is best seen through T-SNE visualizations. In unsupervised spaces, spammers and legitimate users are mixed like oil and water. In the SSDMV latent space, however, they are clearly clustered, making the job of the classifier significantly easier.

Visualization of Latent Space Fig 2: Comparison between standard embedding (top) and SSDMV learned representations (bottom). Red points represent spammers.

Conclusion

SSDMV represents a significant leap in social media security. By moving away from "shallow and supervised" to "deep and semi-supervised," the authors have created a framework that is not only more accurate but also drastically more efficient for real-world deployment where labels are a luxury.

Future Outlook: The next frontier for SSDMV could be its application in adversarial settings where spammers intentionally camouflage their behavior to mimic human multi-view patterns.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Gated Neural Networks or Filter Gates for multi-modal data fusion in social media analysis.
  • What is the origin of the "Ladder Network" for semi-supervised learning, and how have subsequent works adapted its skip-connection architecture for graph-structured data?
  • Explore if the SSDMV framework has been extended to real-time stream processing or applied to cross-platform spammer detection (e.g., across Twitter and Facebook).
Contents
SSDMV: Shielding Social Networks with Deep Multi-View Fusion and Semi-Supervised Learning
1. TL;DR
2. Problem & Motivation: The "Shallow" Gap and Label Scarcity
3. Methodology: The Correlated-Ladder-Network (CLN)
3.1. The Innovation: The Filter Gate
3.2. Unified Optimization
4. Experiments: Doing More with Less
4.1. Robustness to Noise and Sparsity
5. Deep Insights: Visualizing the Latent Space
6. Conclusion