Deep Learning for Social Links: Beyond Simple Statistics with RBMs and DBNs

Deep Learning Approaches for Link Prediction in Social Network Services

2013-01-01
Feng Liu, Bingquan Liu, Chengjie Sun, Ming Liu, Xiaolong Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces deep learning frameworks for link prediction in Social Network Services (SNS) using Restricted Boltzmann Machines (RBM) and Deep Belief Networks (DBN). It presents three distinct approaches: unsupervised link prediction, high-level feature representation, and a generative joint distribution model for link values.

TL;DR

This paper explores how Deep Belief Networks (DBN) and Restricted Boltzmann Machines (RBM) can revolutionize link prediction in Social Network Services (SNS). By moving beyond raw statistical features (like common neighbors) to learned hierarchical representations, the authors achieve higher accuracy in predicting user attitudes (support vs. oppose), even with limited labeled data.

Problem & Motivation

In social platforms like Wikipedia, interactions are more than just "connections"—they carry values (positive or negative). Predicting these values is the Link Prediction problem.

Previously, researchers relied on:

  • Heuristic features: Common neighbors, degrees, etc.
  • Supervised models: Logistic Regression or Decision Trees.

However, these methods suffer from two major flaws:

  1. Label Dependency: They require massive amounts of precisely labeled data.
  2. Feature Limitations: They cannot capture the abstract, non-linear "essence" of social interactions.

The authors' insight was to leverage the high representational power of deep learning to "re-map" these features into a more discriminative latent space.

Methodology: The Three-Pronged Approach

The core of this work lies in the transition from raw data to hierarchical abstraction using RBMs. An RBM is a two-layer stochastic network that learns to represent the distribution of visible data units through hidden units.

RBM and DBN Architectures

The paper proposes three specific implementations:

1. Unsupervised Link Prediction

By using a DBN as a dimensionality reducer (similar to an autoencoder), the model collapses features into a very low-dimensional hidden state. This acts as a clustering mechanism where the "state" of the final hidden unit represents the predicted link value, allowing the model to work even when labels are scarce.

2. Feature Representation Learning

Instead of using raw statistics, the authors feed features through a DBN and use the hidden unit activations as the new feature set. This "feature engineering via deep learning" approach allows even simple classifiers like Logistic Regression to perform better because the input space has been optimized and de-noised.

3. Generative Joint Distribution (The Classifying DBN)

The most sophisticated approach involves a DBN where the top layer's visible units are a "concatenation" of features and labels.

  • Training: The model learns how features and labels "co-exist."
  • Inference: When testing a new link, the model tries all possible labels, calculates the Free Energy () for each, and picks the one with the highest log probability via a Softmax function.

Experiments & SOTA Results

The authors utilized the Wikipedia Requests for Adminship (RfA) dataset. The task was to predict whether a "vote" between users was supportive or opposing based on 28 structural features.

Feature Enhancement Results

Using represented features significantly boosted the baseline performance of standard Logistic Regression models:

DBN StructureAUCPrecision
Original Features (Baseline)0.86179.7%
DBN (28x28) + LR0.88681.8%
DBN (28x56) + LR0.89982.1%

Deep Classifier Success

The Three-RBM DBN model demonstrated that deeper architectures (within reason) provide better abstraction for the link prediction task.

Performance Comparison Table

The model reached a peak Precision of 84.6%, proving that the joint distribution modeling via RBMs is more robust than traditional discriminative models.

Deep Insights & Conclusion

The primary takeaway is the effectiveness of latent representation. Even though the input layer only had 28 features, passing them through hidden RBM layers allowed the model to discover "latent social signatures" that are not visible in raw counts of common neighbors.

Limitations: The authors noted that adding too many layers did not yield further improvements. This is likely due to the low dimensionality of the input; for small feature sets, "over-abstracting" leads to information loss rather than gain.

Future Outlook: This work paves the way for integrating more complex features (like the text content of social comments) into the RBM framework, potentially allowing for a unified model of structural and semantic social dynamics.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Graph Neural Networks (GNNs) or Variational Autoencoders (VAEs) to signed link prediction in social networks specifically for the Wikipedia RfA dataset.
  • Which paper first established the Contrastive Divergence (CD) learning procedure for RBMs, and how has its efficiency evolved compared to the methods used in this study?
  • Explore how the joint distribution modeling approach used here for link prediction has been extended to multi-modal link prediction tasks involving both text and graph structures.
Contents
Deep Learning for Social Links: Beyond Simple Statistics with RBMs and DBNs
1. TL;DR
2. Problem & Motivation
3. Methodology: The Three-Pronged Approach
3.1. 1. Unsupervised Link Prediction
3.2. 2. Feature Representation Learning
3.3. 3. Generative Joint Distribution (The Classifying DBN)
4. Experiments & SOTA Results
4.1. Feature Enhancement Results
4.2. Deep Classifier Success
5. Deep Insights & Conclusion