Beyond Follower Graphs: Clustering Social Media Users via Probabilistic Topic Modeling

Clustering Users in Micro Blogging Social Networks Using Probabilistic Topic Modeling - A Framework

2012-06-01
Hossein Dolatabadi, Lay-Ki Soon, Mahdi Negahi Shirazi, Mohammad Mohammadi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a framework for clustering users in micro-blogging social networks like Twitter using probabilistic topic modeling. By implementing a modified Latent Dirichlet Allocation (Twitter-LDA), the methodology extracts latent interests from short-form user posts to group individuals with similar topical distributions.

TL;DR

This research introduces a methodology to categorize micro-blogging users not by who they follow (structural links), but by what they talk about (topical content). By refining Latent Dirichlet Allocation (LDA) for the constraints of Twitter, the authors provide a framework to group users based on their latent interests, overcoming the limitations of short, noisy text.

The Structural Fallacy: Why Links Aren't Enough

In platforms like Twitter, the "Following" mechanism is the standard for grouping. However, this creates a Structural Fallacy: following a journalist or a celebrity does not necessarily mean the user shares the same interests or produces similar content.

Existing methods struggle with:

  • Data Sparsity: Tweets are restricted (historically 140 characters), providing very little context for standard NLP.
  • Linguistic Noise: The prevalence of abbreviations, slang, and misspellings creates a "dirty" dataset that breaks traditional Latent Semantic Indexing.
  • Asymmetric Relationships: Unlike Facebook's "friendship," Twitter's one-way following doesn't guarantee mutual interest.

Methodology: The Twitter-LDA Pipeline

The authors propose a robust pipeline that moves from raw RSS feeds to a "Topic Probability Table."

1. Robust Pre-processing

To tackle the noise, the framework employs:

  • Stop-word Removal: Filtering 647 common English terms.
  • Dictionary Normalization: Using a 45,000-word vocabulary to correct misspellings and expand abbreviations—a critical step for maintaining the integrity of the word-topic distribution.

2. Twitter-LDA: Adapted Inference

Standard LDA assumes a document is a mixture of multiple topics. In contrast, Twitter-LDA (as utilized by the authors) operates on the intuition that a single tweet is usually about one specific thing. This constraint significantly reduces the noise in the probabilistic distribution.

System Algorithm Workflow Note: The image above illustrates the RSS capture process, the precursor to the LDA inference.

Experiments and Results

The study focused on a specific cohort: Malaysian journalists and their followers (229 core users).

  • Scale: The system extracted the maximum allowable 3,200 tweets per user via the Twitter API.
  • Granularity: The LDA was configured to find 20 latent topics across the dataset.
  • Outcome: The result is a User-Topic Matrix. If User A and User B both have a high probability density for "Politics" and "Technology," they are clustered together, regardless of whether they follow each other.

Critical Insight: The Value of "Topic Sensors"

The paper highlights a fascinating perspective: viewing every user as a natural event sensor. When topic modeling is applied in real-time, clusters of users shifting their focus to a specific topic (e.g., "Earthquake") can serve as a faster detection system than traditional media.

Deep Takeaway

For commercial and governmental entities, this framework moves beyond "Demographics" and into "Psychographics." It allows for:

  1. Market Segmentation: Micro-targeting based on actual discourse.
  2. Harm Prevention: Monitoring clusters discussing harmful activities.
  3. Platform Optimization: Helping social media providers suggest relevant content to keep users engaged.

Limitations & Future Work

While robust, the current framework is primarily focused on English. The evolution of this work would naturally involve Cross-lingual Topic Modeling to handle the multilingual nature of regions like Malaysia. Furthermore, integrating temporal analysis (how interests change over time) would add a dynamic layer to the current static clusters.

Find Similar Papers

Try Our Examples

  • Search for recent studies that combine structural graph features with Twitter-LDA for hybrid social network clustering.
  • What are the benchmark datasets currently used to evaluate the accuracy of user interest discovery in micro-blogging platforms?
  • How do modern transformer-based models (like BERTopic) compare to Twitter-LDA for short-text topic modeling in 2024-2025?
Contents
Beyond Follower Graphs: Clustering Social Media Users via Probabilistic Topic Modeling
1. TL;DR
2. The Structural Fallacy: Why Links Aren't Enough
3. Methodology: The Twitter-LDA Pipeline
3.1. 1. Robust Pre-processing
3.2. 2. Twitter-LDA: Adapted Inference
4. Experiments and Results
5. Critical Insight: The Value of "Topic Sensors"
5.1. Deep Takeaway
6. Limitations & Future Work