Probabilistic Lexicon Expansion: Bridging the Gap in Social Media Slang
A probalistic approach to automatically extract new words from social media
This paper introduces a probabilistic framework for the automatic extraction of new keywords from short-form social media content (specifically Twitter). By leveraging Bayes Rule and Jaro-Winkler distance, the method identifies links between unknown terms and existing thematic dictionaries to dynamically update lexicons for community detection.
TL;DR
Social media moves faster than any dictionary. To solve this, the authors propose a probabilistic framework that automatically identifies "new" words (neologisms) from short-form messages like Tweets and maps them to specific communities (e.g., Education, Bullying, Terrorism). By using conditional probability and graph-based link pruning, they turn static lexicons into dynamic graphs that evolve alongside user behavior.
Context & Motivation
In the world of Natural Language Processing (NLP), micro-blogs are a nightmare. Standard tools built for research papers or news articles fail because:
- Length Constraints: A 140-character tweet doesn't have enough "surface area" for traditional TF-IDF or positional features.
- Vocabulary Volatility: Slang, abbreviations, and coded language (especially in malicious communities) appear and change almost daily.
The authors argue that identifying new keywords is not just about detection; it's about association. If a new word frequently appears alongside known "Bullying" terms, the system should learn to treat that new word as part of the bullying lexicon.
Methodology: The Bayes-Jaro Pipeline
The proposed framework follows a sophisticated four-stage process:
1. Linguistic Pre-processing
Using the Stanford POS Tagger, the system strips away the "noise" (prepositions, adverbs) and focuses on the "signal" (nouns, adjectives, verbs). This is followed by stemming to reduce words to their base forms (e.g., "fishing" to "fish").
2. Similarity Matching with Jaro-Winkler
How do you know if a word is truly "new"? The system compares tokens against existing dictionaries using Jaro-Winkler Distance. This is particularly effective for short strings because it gives more weight to the prefix of the string, making it resilient to slight spelling variations common on Twitter.
3. Probabilistic Graph Construction
This is the core "intelligence" of the paper. Instead of just counting frequencies, the authors build a directed graph where:
- Nodes: Both existing () and new keywords ().
- Edges: Determined by conditional probability —the likelihood of a new word appearing given a known keyword.

4. Chi-Square Link Pruning
Not every pairing is meaningful. To separate true associations from random noise, the authors apply a Chi-Square test. If the dependency between two words is statistically significant (p < 0.05), the link is retained; otherwise, it is pruned.
Experiments & Real-World Results
The researchers tested their approach on 1,000 live tweets, generating a pool of roughly 2,900 unique keywords.
Key Benchmarks:
- Jaro-Winkler Accuracy: 92.5%. The system is excellent at distinguishing known terms from unknown ones.
- Domain Assignment Accuracy: 74.4%. While solid, this reflects the difficulty of the task. Errors often occurred when a word was used across multiple domains, making it a "general" word rather than a domain-specific one.
In the figure above, the word "fat" was statistically linked to "hate" and "mad," allowing the "Bullying" dictionary to be automatically updated with this "new" context.
Critical Insight & Future Outlook
The beauty of this approach is its Domain Independence. Whether you are tracking sports trends or monitoring online radicalization, the underlying math remains the same.
Limitations to Consider:
- Data Sparsity: The 74.4% accuracy suggests that 1,000 tweets are not enough. Probabilistic models thrive on Big Data; a larger corpus would likely reduce the error rate.
- Context Shifting: A word might mean something in the "Sports" community today and something entirely different in a "Political" community tomorrow.
The Takeaway: This research provides a vital blueprint for building living lexicons. In the future, combining this probabilistic approach with deep learning (like Transformer-based embeddings) could create even more robust systems for real-time social media intelligence.
