intJNMF: Solving Twitter's Sparsity Problem by Mixing Text with Social DNA
What and With Whom? Identifying Topics in Twitter Through Both Interactions and Text
This paper introduces intJNMF, a novel topic derivation method for Twitter that combines tweet content similarity with social interaction features (mentions, replies, and retweets). By utilizing a two-step Non-negative Matrix Factorization (NMF) process, it achieves a significant SOTA advancement, outperforming traditional methods like LDA and standard NMF by over 30% in clustering accuracy.
TL;DR
Deriving topics from Twitter is notoriously difficult because tweets are too short for traditional AI to find common patterns. The paper "What and With Whom?" introduces intJNMF, a technique that stops looking just at what people say and starts looking at who they are talking to. By combining mentions, replies, and retweets into a mathematical "relationship matrix," the authors achieved a 30%+ boost in topic accuracy over standard models.
The "Zero-Overlap" Nightmare
If User A tweets "The new Senate is exciting!" and User B replies "True, the census in Australia is a mess," a traditional model like Latent Dirichlet Allocation (LDA) will likely fail to link them. Why? Because they share zero keywords.
In the world of Twitter, data is "sparse." The standard Tweet-Term Matrix (mapping tweets to words) is usually over 99.9% empty. This makes conventional topic modeling a shot in the dark. The authors realized that while words might be missing, interactions (replies, retweets) are "hard signals" that two posts belong together.
Methodology: The Two-Step Factorization
The core innovation of intJNMF is its two-step dance using Non-negative Matrix Factorization (NMF).
Step 1: Building the Social Map
Instead of starting with words, the authors build a Tweet-Relationship Matrix (A). This matrix calculates the connection between Tweet and Tweet based on:
- Action Similarity: Are they part of a reply/retweet chain? (Strongest signal).
- People Similarity: Do they mention the same users?
- Content Similarity: Traditional cosine similarity (the fallback).

Step 2: From Clusters to Keywords
Most NMF approaches try to solve for clusters and words simultaneously. intJNMF decouples them. It first factorizes the dense relationship matrix to find "hidden clusters" (latent factors). Then, it "freezes" these clusters and uses them to extract the most relevant keywords from the sparse text data. This "Joint" approach ensures that even if a tweet has unique words, it gets pulled into the correct topic by its social gravity.
Experimental Showdown
The authors tested their method against the TREC2014 (standard benchmark) and a custom tweetMarch dataset.
Performance Boost
The results were dramatic. In metrics like Purity and NMI (measures of how "clean" the topic clusters are), intJNMF outperformed the standard LDA and NMF models by significant margins. In the TREC2014 dataset, intJNMF's NMI score was 0.367, compared to a measly 0.058 for standard NMF.

Why it Works: Interaction Impact
The researchers performed an Ablation Study to see which feature helped most. While content similarity is the most common feature (connecting the most tweets), the "Action" features (replies/retweets) had a 99% accuracy rate in predicting a shared topic. By combining all three, the model reached a synergy that single-feature models couldn't touch.
Human Readability
A key takeaway for practitioners is the "readability" of the keywords. When the model used social context, the keywords it produced for a topic like "Politics" were highly coherent (e.g., liberal, obama, government), whereas standard NMF produced noise (e.g., gain, high).
Conclusion & Future Outlook
The paper effectively proves that context is king. In the era of LLMs and massive embeddings, this work reminds us that structured metadata (social graphs) provides an structural "anchor" that text alone cannot provide.
Limitations: The model is currently batch-processed. As Twitter moves in real-time, the next frontier for this research is an incremental model that can update topics as quickly as a "trending" hashtag appears.
Takeaway: If you are building a recommendation engine or a news aggregator, don't just analyze the text. Look at the interactions—they are the blueprint of the conversation.
