UATM: Deciphering User Intentions in Sparse Social Networks via Aggregated Topic Modeling

A user-based aggregation topic model for understanding user’s preference and intention in social network

2020-07-06
Lei Shi, Guangjia Song, Gang Cheng, Xia Liu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the User-based Aggregation Topic Model (UATM), a novel framework designed to mine user preferences and intentions in social networks. By aggregating content from both users and their followees and incorporating an RNN-based weight prior, UATM achieves SOTA performance in short-text topic coherence and user clustering.

Executive Summary

TL;DR

Understanding what a user wants from a 140-character tweet is notoriously difficult due to "data sparsity." This paper presents UATM (User-based Aggregation Topic Model), which survives the sparsity desert by aggregating a user's posts with those of their followees and injecting neural semantic priors from an Elman RNN. The result is a robust model that identifies coherent interests and clusters users with high precision without ever touching private search logs.

Background Positioning

Unlike traditional LDA which views documents in isolation, UATM is a SOTA enhancement for short-text analysis. It sits at the intersection of probabilistic graphical models and recurrent neural networks, moving beyond simple word counts to capture the "intent" hidden in social structures.

The Sparsity Wall: Why Social Mining is Hard

Social network data presents a unique paradox: it is massive in volume but microscopically thin in context. Prior works failed primarily because:

  1. Contextual Sparsity: A single tweet lacks enough word co-occurrences for traditional LDA to form stable clusters.
  2. Privacy Barriers: Commercial systems use click-stream data, but for researchers, this data is behind a "privacy wall."
  3. Noise Overload: Social media is 80% background noise (common words) and 20% intent. Distinguishing the two is a major hurdle.

Methodology: The Core Innovations

UATM attacks the problem through a three-layered architecture.

1. The RNN-IDF Weight Prior

Instead of assuming all words are independent, the authors use an Elman RNN to learn the quantifiable relationship between words. This neural prior is combined with Inverse Document Frequency (IDF) to weight the importance of word pairs. If the RNN says "Apple" and "iPhone" are highly related in your corpus, UATM gives them a higher prior probability of sharing a topic.

2. User-Followee Aggregation

UATM assumes that "you are who you follow." By modeling the topic distribution of a user's followees () alongside the user's own content (), the model gains a massive amount of auxiliary context to smooth out the sparsity of a single user's timeline.

Overall Framework Figure: The graphical representation of UATM shows the interplay between user interests (), followee interests (), and the word-pair generation process.

3. Biterm Modeling with Background Switch

Following the success of Biterm Topic Models, UATM models word-pairs () globally. Crucially, it uses a Bernoulli switch () to decide: is this pair just "background noise" (common language) or a "topical signal"?

Experiments and Results

The authors validated UATM on a massive Sina Weibo dataset.

Qualitative Coherence

In a comparison of topic words for the "Xinjiang Earthquake," UATM generated highly specific terms like "Focus," "Depth," and "Yutian County," whereas LDA and other baselines were polluted with noise like "Tourism" or "Beef."

Quantitative Edge

Across metrics like PMI-score (coherence) and NMI (clustering quality), UATM consistently held the lead.

Clustering Results Figure: Performance comparison across Purity, NMI, and ARI shows UATM (red line) consistently outperforming baselines as the number of topics increases.

Key finding: The optimal weight for followee information was found to be around , suggesting that your social circle is actually a stronger predictor of your interests than your individual (and often sparse) posts.

Critical Insight & Future Outlook

Takeaway: The success of UATM proves that "structural context" (who you follow) is the ultimate antidote to "content sparsity" (short tweets). By mathematically blending social graphs with neural word embeddings, we can build recommendation engines that respect privacy while remaining highly accurate.

Limitations: The model is currently static. Social interests drift over time, and location matters (spatial-temporal context). Future Work: The logical next step is a Dynamic UATM that incorporates time-slices and GPS data to track how intentions evolve during live events.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) instead of simple aggregation to model followee relationships in short-text topic modeling.
  • What are the seminal papers that first introduced the Biterm Topic Model (BTM), and how have subsequent works integrated deep learning priors into its Dirichlet structure?
  • Find research that applies the UATM framework or similar user-aggregation logic to multi-modal social media data, such as combined image-text posts on platforms like Instagram.
Contents
UATM: Deciphering User Intentions in Sparse Social Networks via Aggregated Topic Modeling
1. Executive Summary
1.1. TL;DR
1.2. Background Positioning
2. The Sparsity Wall: Why Social Mining is Hard
3. Methodology: The Core Innovations
3.1. 1. The RNN-IDF Weight Prior
3.2. 2. User-Followee Aggregation
3.3. 3. Biterm Modeling with Background Switch
4. Experiments and Results
4.1. Qualitative Coherence
4.2. Quantitative Edge
5. Critical Insight & Future Outlook