KETNEW: Enhancing Twitter Keyword Extraction via Hybrid Graph Weighting
A Graph based Keyword Extraction from Twitter using Node and Edge Weight
The paper introduces KETNEW (Keyword Extraction from Tweets using Node and Edge Weight), an unsupervised graph-based hybrid approach for keyword extraction from social media. It ranks keywords by computing Node Importance (NI) scores derived from structural graph properties, statistical linguistics, and positional features.
TL;DR
Twitter's brevity makes traditional keyword extraction a nightmare. KETNEW (Keyword Extraction from Tweets using Node and Edge Weight) solves this by moving beyond simple word counts. It treats a collection of tweets as a complex network, assigning importance based on where a word appears, how often it’s tweeted, and how tightly its "friends" (neighboring words) are connected.
The Problem: The Chaos of Micro-blogging
Traditional Information Retrieval (IR) models like TF-IDF treat documents as "bags of words," ignoring the sequence and structure. In a 280-character tweet, this is a fatal flaw. Furthermore, classic graph models like TextRank often treat all edges equally.
The authors identify three main gaps in current SOTA:
- Contextual Ignorance: Ignoring whether a word appears at the start (indicative of topic) or end of a tweet.
- Structural Simplicity: Using only degree centrality instead of more nuanced metrics like the Clustering Coefficient.
- Noise Sensitivity: Failing to account for the repetitive nature of retweets and the clutter of URLs/hashtags.
Methodology: The NI (Node Importance) Score
The core innovation of KETNEW is its scoring mechanism. It builds an undirected weighted graph and calculates weights using a "Hybrid Insight" approach.
1. Node Weighting (The "Identity" of a Word)
Instead of just frequency, node weight is a composite of:
- Tweet Frequency: How many unique tweets contain the word.
- Positional Weight: Bits specifically for appearing at the
firstorlastposition. - Clustering Coefficient: A measure of how much the word's neighbors are connected to each other, highlighting "clique-ish" topical terms.
2. Edge Weighting (The "Relationship")
Edges aren't just binary. They are weighted by:
- Co-occurrence Frequency: How often two words appear together.
- Shared Neighbors: A normalization of local connectivity overlap.
(Note: Refer to the paper's Algorithm 1 for the iterative NI score calculation)
Experiments & Results
The researchers tested KETNEW against five major news topics (e.g., NASA programs, Google-Motorola deals) using the First Story Detection (FSD) dataset.
SOTA Comparison
The results were clear: KETNEW and its Eigenvector variation dominated.
- TF-IDF struggled significantly, often returning a Precision of 0.0 in specific news categories because it couldn't handle the sparse nature of the data.
- TextRank performed moderately (F-measure ~0.55) but lacked the precision of KETNEW’s multi-feature scoring.
Fig. 1. Average Performance Analysis: KETNEW shows a visible lead over TF-IDF and traditional Graph Centrality measures.
Critical Analysis & Conclusion
Why it works
KETNEW succeeds because it acknowledges the Linguistic vs. Structural trade-off. By incorporating the "Position" of a word, it captures a linguistic truth (important stuff usually comes first), and by using "Clustering Coefficients," it captures a structural truth (keywords form the core of topical clusters).
Limitations
- Computational Complexity: Calculating clustering coefficients and shared neighbors for very large graphs (millions of nodes) can be expensive compared to simple TF-IDF.
- Language Dependency: While the authors claim portability, the "Position" feature might vary across languages with significantly different syntax (e.g., SOV vs SVO).
Final Takeaway
KETNEW proves that even in the world of Deep Learning, unsupervised graph algorithms remain powerful and interpretable. For developers building real-time trend detectors or summarizers, KETNEW provides a lightweight yet efficient roadmap.
