STMP: Overcoming Short-Text Sparsity in Sentiment-Aware Topic Modeling

A Sentiment and Topic Model with Timeslice, User and Hashtag for Posts on Social Media

2017-01-01
Kang Xu, Junheng Huang, Tianxing Wu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Sentiment Topic Model for Posts (STMP), a novel joint sentiment-topic model specifically designed for social media short texts. By integrating structure information—including timeslices, user identity, and a mined "potential hashtag" expansion—the model aggregates sparse posts into dense pseudo-documents to achieve state-of-the-art results in sentiment-aware topic extraction.

TL;DR

Twitter and other micro-blogging platforms are goldmines for public opinion, but their 140-character limit breaks traditional NLP models. The Sentiment Topic Model for Posts (STMP) fixes this by aggregating sparse tweets into "pseudo-documents" using Users, Timeslices, and Hashtags as anchors. By expanding explicit hashtags with semantically related "potential hashtags," STMP effectively "thickens" the data, allowing it to outperform established models like JST and ASUM.

The Sparsity Crisis in Social Media

Most topic models (like LDA) were built for long-form documents where words have plenty of neighbors to establish context. Social media posts are the opposite: they are sparse and informal. If two tweets discuss the same topic but use different words, traditional models often fail to link them.

The authors identify a missed opportunity: Social media isn't just text; it’s a web of structure information. Every post comes with a timestamp, a user ID, and often, a hashtag. STMP's core intuition is that these "metadata anchors" can give us the context that the text alone lacks.

Methodology: The Power of Metadata Anchors

The signature innovation of STMP is how it organizes the generative process around three pillars:

  1. Temporal Topics (Timeslices): Capturing "bursty" events or trends currently happening.
  2. Stable Topics (Users): Capturing long-term personal interests or professional focus.
  3. Hashtag Topics: Using tags as explicit topic labels.

Mining Potential Hashtags

A unique pre-processing step involves expanding explicit hashtags (e.g., #iPhone) into potential hashtags (e.g., #Apple, #Smartphone). The authors use Word2Vec and cosine similarity to find these related tags, ensuring that even if a user forgets to tag a post, the model can infer the thematic connection based on word embeddings.

Model Architecture and Topic Coherence Results Figure 1: Comparison of STMP against Baselines on Topic Coherence and Precision.

Experiments: Superior Quality

Testing on the Twitter7 dataset (specifically focused on electronic products), STMP was compared against traditional models (JST, ASUM) and an earlier metadata-driven model (TUS-LDA).

  • Topic Coherence: STMP consistently produced more human-interpretable topics.
  • Precision@20: The model showed a higher density of "correct" topic-related words in the top 20 results compared to all baselines.

Qualitative Look: The Sentiment-Topic Matrix

The model successfully separated positive and negative sentiment within specific technical domains:

Topic (Positive)Key WordsTopic (Negative)Key Words
Cameradigit, canon, nikon, slrPrinterink, cartridge, laser, slow
Musicipod, song, music, lovePhoneproblem, security, strange, risk

Experimental Results Table Table 1: Top words for positive and negative topics extracted by STMP.

Critical Insight & Conclusion

The success of STMP proves that metadata is not just noise; it is a feature. By leveraging the "Who, When, and What (Hashtag)" of a post, the model compensates for the lack of linguistic depth in short texts.

Future Outlook: While STMP uses Word2Vec for hashtag expansion, the authors suggest the next leap will be fully integrating word embeddings into the Dirichlet distribution itself. For practitioners, the takeaway is clear: when dealing with short-form content, don't just look at the words—look at the context in which they were born.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) or Contrastive Learning to solve the short-text context sparsity problem in social media topic modeling.
  • Which paper originally proposed the concept of "pseudo-document aggregation" for microblogs, and how has the selection of aggregation keys evolved since then?
  • Are there any studies that apply the STMP framework or similar metadata-driven topic models to multi-modal social media posts (e.g., combining text with image features)?
Contents
STMP: Overcoming Short-Text Sparsity in Sentiment-Aware Topic Modeling
1. TL;DR
2. The Sparsity Crisis in Social Media
3. Methodology: The Power of Metadata Anchors
3.1. Mining Potential Hashtags
4. Experiments: Superior Quality
4.1. Qualitative Look: The Sentiment-Topic Matrix
5. Critical Insight & Conclusion