Selective Listening: Advancing Informative Tweet Detection via Weighted Naive Bayes
Informative vs. Non-informative Short Message Detection in Social Networks
This paper introduces a Weighted Binary Multinomial Naive Bayes (BMNB) variation designed to classify tweets into "informative" (public interest/news) and "non-informative" (personal/noise) categories. By incorporating specialized feature weighting for hashtags and user mentions alongside a data-driven prior distribution, the method significantly enhances the retrieval of useful social media content.
TL;DR
In the deluge of social media, separating the signal (news, trends, public events) from the noise (personal chatter, mood updates) is a critical preprocessing step. This paper proposes a Weighted Binary Multinomial Naive Bayes (BMNB) model that treats hashtags and user mentions as "premium" features. By replacing generic smoothing with a data-driven prior, the model achieves a recall of up to 94%, ensuring that almost no informative content is lost while effectively downsizing massive datasets.
Problem & Motivation: The Sparsity Trap
Why is Twitter data so hard to clean? The authors identify three primary "sparsity traps":
- Short Length: Traditional Topic Models like LDA fail because there isn't enough word co-occurrence in 140-280 characters.
- Frequency Paradox: In short messages, TF-IDF weights are often meaningless because terms rarely repeat within a single document.
- Linguistic Chaos: Slang, typos, and multilingual tokens make standardized dictionary indexing highly inefficient.
The core motivation here isn't just classification—it's data reduction. If we can accurately discard "personal noise" (non-informative tweets) without losing public updates, we drastically reduce dimensionality for downstream tasks like recommendation engines or trend detection.
Methodology: Beyond Uniform Smoothing
The authors' "Secret Sauce" involves two key surgical modifications to the standard BMNB model:
1. The Physics of Weighting
Instead of treating every token equally, the model categorizes tokens into three buckets: Words, Hashtags, and Mentioned Users. Through empirical testing, they discovered that user mentions are the strongest indicators of informativeness (often pointing to journalists, celebrities, or organizations). The conditional probability is modified as: Where is the weight assigned to the specific feature type.
2. Prior Distribution vs. Laplace
Standard Naive Bayes uses Laplace (+1) smoothing, which assumes a Uniform Prior. The authors argue this is suboptimal. Instead, they sample 10% of the training data to estimate a Dirichlet Prior that reflects the actual distribution of terms in the specific corpus, making the model much more "aware" of the environment it is operating in.

Experiments & Results: High Recall, Low Noise
The authors tested their approach on two independent datasets (Dataset A: ~20k tweets; Dataset B: ~10k tweets).
Key Findings:
- The Weight of a User: The best performance was achieved when mentioned users were weighted significantly higher (e.g., ) than standard words ().
- Recall is King: Since the goal is filtering, the authors prioritized scores (weighting recall over precision). Their model maintained a recall of ~92-94%, meaning it barely missed any "useful" tweets.
- Efficiency: The model successfully identified and labeled over 6,000 tweets in Dataset A as non-informative. For a production system, this means a 30% reduction in data volume with negligible information loss.
Fig. 1. TP comparison: The weighted model (green) consistently captures more informative messages than the baseline BMNB.
Critical Analysis & Conclusion
The value of this work lies in its simplicity and interpretability. While modern LLMs could solve this classification task, a Weighted BMNB is computationally "cheap" and can run in real-time on massive firehose streams without GPU clusters.
Limitations:
- Manual Weighting: The weights were determined manually. Future iterations could benefit from automated hyperparameter optimization or attention-based weighting.
- Context Blindness: BMNB still treats words as a "bag," potentially missing nuanced sarcasm or context-dependent informativeness.
Takeaway: This paper proves that even "old" algorithms like Naive Bayes can be highly competitive if we inject domain-specific structural knowledge (like the inherent value of a Hashtag) into the mathematical prior.
