intJNMF: Solving Twitter's Sparsity Problem by Mixing Text with Social DNA

What and With Whom? Identifying Topics in Twitter Through Both Interactions and Text

2017-04-24
Robertus Nugroho, Jian Yang, Weiliang Zhao, Cécile Paris, Surya Nepal
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces intJNMF, a novel topic derivation method for Twitter that combines tweet content similarity with social interaction features (mentions, replies, and retweets). By utilizing a two-step Non-negative Matrix Factorization (NMF) process, it achieves a significant SOTA advancement, outperforming traditional methods like LDA and standard NMF by over 30% in clustering accuracy.

TL;DR

Deriving topics from Twitter is notoriously difficult because tweets are too short for traditional AI to find common patterns. The paper "What and With Whom?" introduces intJNMF, a technique that stops looking just at what people say and starts looking at who they are talking to. By combining mentions, replies, and retweets into a mathematical "relationship matrix," the authors achieved a 30%+ boost in topic accuracy over standard models.

The "Zero-Overlap" Nightmare

If User A tweets "The new Senate is exciting!" and User B replies "True, the census in Australia is a mess," a traditional model like Latent Dirichlet Allocation (LDA) will likely fail to link them. Why? Because they share zero keywords.

In the world of Twitter, data is "sparse." The standard Tweet-Term Matrix (mapping tweets to words) is usually over 99.9% empty. This makes conventional topic modeling a shot in the dark. The authors realized that while words might be missing, interactions (replies, retweets) are "hard signals" that two posts belong together.

Methodology: The Two-Step Factorization

The core innovation of intJNMF is its two-step dance using Non-negative Matrix Factorization (NMF).

Step 1: Building the Social Map

Instead of starting with words, the authors build a Tweet-Relationship Matrix (A). This matrix calculates the connection between Tweet and Tweet based on:

  • Action Similarity: Are they part of a reply/retweet chain? (Strongest signal).
  • People Similarity: Do they mention the same users?
  • Content Similarity: Traditional cosine similarity (the fallback).

Tweet Relationship Modeling

Step 2: From Clusters to Keywords

Most NMF approaches try to solve for clusters and words simultaneously. intJNMF decouples them. It first factorizes the dense relationship matrix to find "hidden clusters" (latent factors). Then, it "freezes" these clusters and uses them to extract the most relevant keywords from the sparse text data. This "Joint" approach ensures that even if a tweet has unique words, it gets pulled into the correct topic by its social gravity.

Experimental Showdown

The authors tested their method against the TREC2014 (standard benchmark) and a custom tweetMarch dataset.

Performance Boost

The results were dramatic. In metrics like Purity and NMI (measures of how "clean" the topic clusters are), intJNMF outperformed the standard LDA and NMF models by significant margins. In the TREC2014 dataset, intJNMF's NMI score was 0.367, compared to a measly 0.058 for standard NMF.

Comparison of Clustering Results

Why it Works: Interaction Impact

The researchers performed an Ablation Study to see which feature helped most. While content similarity is the most common feature (connecting the most tweets), the "Action" features (replies/retweets) had a 99% accuracy rate in predicting a shared topic. By combining all three, the model reached a synergy that single-feature models couldn't touch.

Human Readability

A key takeaway for practitioners is the "readability" of the keywords. When the model used social context, the keywords it produced for a topic like "Politics" were highly coherent (e.g., liberal, obama, government), whereas standard NMF produced noise (e.g., gain, high).

Conclusion & Future Outlook

The paper effectively proves that context is king. In the era of LLMs and massive embeddings, this work reminds us that structured metadata (social graphs) provides an structural "anchor" that text alone cannot provide.

Limitations: The model is currently batch-processed. As Twitter moves in real-time, the next frontier for this research is an incremental model that can update topics as quickly as a "trending" hashtag appears.

Takeaway: If you are building a recommendation engine or a news aggregator, don't just analyze the text. Look at the interactions—they are the blueprint of the conversation.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Graph Neural Networks (GNNs) instead of NMF to integrate social interactions and text for Twitter topic modeling.
  • Which paper first introduced the concept of "Short Text Topic Modeling" (STTM) as a distinct subfield, and how does the Joint-NMF approach address the sparsity issues defined there?
  • Find research that applies the intJNMF methodology or similar interaction-based clustering to multi-modal social media platforms like Instagram or TikTok.
Contents
intJNMF: Solving Twitter's Sparsity Problem by Mixing Text with Social DNA
1. TL;DR
2. The "Zero-Overlap" Nightmare
3. Methodology: The Two-Step Factorization
3.1. Step 1: Building the Social Map
3.2. Step 2: From Clusters to Keywords
4. Experimental Showdown
4.1. Performance Boost
4.2. Why it Works: Interaction Impact
5. Human Readability
6. Conclusion & Future Outlook