STMP: Overcoming Short-Text Sparsity in Sentiment-Aware Topic Modeling
A Sentiment and Topic Model with Timeslice, User and Hashtag for Posts on Social Media
The paper introduces the Sentiment Topic Model for Posts (STMP), a novel joint sentiment-topic model specifically designed for social media short texts. By integrating structure information—including timeslices, user identity, and a mined "potential hashtag" expansion—the model aggregates sparse posts into dense pseudo-documents to achieve state-of-the-art results in sentiment-aware topic extraction.
TL;DR
Twitter and other micro-blogging platforms are goldmines for public opinion, but their 140-character limit breaks traditional NLP models. The Sentiment Topic Model for Posts (STMP) fixes this by aggregating sparse tweets into "pseudo-documents" using Users, Timeslices, and Hashtags as anchors. By expanding explicit hashtags with semantically related "potential hashtags," STMP effectively "thickens" the data, allowing it to outperform established models like JST and ASUM.
The Sparsity Crisis in Social Media
Most topic models (like LDA) were built for long-form documents where words have plenty of neighbors to establish context. Social media posts are the opposite: they are sparse and informal. If two tweets discuss the same topic but use different words, traditional models often fail to link them.
The authors identify a missed opportunity: Social media isn't just text; it’s a web of structure information. Every post comes with a timestamp, a user ID, and often, a hashtag. STMP's core intuition is that these "metadata anchors" can give us the context that the text alone lacks.
Methodology: The Power of Metadata Anchors
The signature innovation of STMP is how it organizes the generative process around three pillars:
- Temporal Topics (Timeslices): Capturing "bursty" events or trends currently happening.
- Stable Topics (Users): Capturing long-term personal interests or professional focus.
- Hashtag Topics: Using tags as explicit topic labels.
Mining Potential Hashtags
A unique pre-processing step involves expanding explicit hashtags (e.g., #iPhone) into potential hashtags (e.g., #Apple, #Smartphone). The authors use Word2Vec and cosine similarity to find these related tags, ensuring that even if a user forgets to tag a post, the model can infer the thematic connection based on word embeddings.
Figure 1: Comparison of STMP against Baselines on Topic Coherence and Precision.
Experiments: Superior Quality
Testing on the Twitter7 dataset (specifically focused on electronic products), STMP was compared against traditional models (JST, ASUM) and an earlier metadata-driven model (TUS-LDA).
- Topic Coherence: STMP consistently produced more human-interpretable topics.
- Precision@20: The model showed a higher density of "correct" topic-related words in the top 20 results compared to all baselines.
Qualitative Look: The Sentiment-Topic Matrix
The model successfully separated positive and negative sentiment within specific technical domains:
| Topic (Positive) | Key Words | Topic (Negative) | Key Words |
|---|---|---|---|
| Camera | digit, canon, nikon, slr | Printer | ink, cartridge, laser, slow |
| Music | ipod, song, music, love | Phone | problem, security, strange, risk |
Table 1: Top words for positive and negative topics extracted by STMP.
Critical Insight & Conclusion
The success of STMP proves that metadata is not just noise; it is a feature. By leveraging the "Who, When, and What (Hashtag)" of a post, the model compensates for the lack of linguistic depth in short texts.
Future Outlook: While STMP uses Word2Vec for hashtag expansion, the authors suggest the next leap will be fully integrating word embeddings into the Dirichlet distribution itself. For practitioners, the takeaway is clear: when dealing with short-form content, don't just look at the words—look at the context in which they were born.
