PWS-PSS: Decoding Sentiment via Semantic Vector Geometry
15512_Unlock big data emotions Weighted word embeddings for sentiment classification.
This paper introduces a sentiment analysis framework utilizing word embeddings and a point-wise scoring system. By calculating the "Positive-Word-Score" (PWS) for individual words and aggregating them into a "Phrase-Sentiment-Score" (PSS), the method achieves up to 78.52% accuracy on domain-specific tweets across Sports, Politics, and Services.
TL;DR
This research presents a scalable framework for sentiment analysis that translates word embedding distances into polarity scores. By identifying the geometric position of a word relative to "Positive" and "Negative" clusters in a vector space, the system calculates a Phrase-Sentiment-Score (PSS). It achieves a peak accuracy of 78.52% on real-world Twitter data, proving that even simple linear combinations of semantic weights can compete in complex domain-specific tasks.
Motivation: The Context Gap in Sentiment
Why is sentiment analysis still challenging despite the rise of LLMs? The answer lies in domain specificity and lexical evolution. Standard dictionaries (e.g., SentiWordNet) often fail to capture how a word like "epic" or "rush" might shift meaning in sports versus services contexts. The authors hypothesize that if word embeddings (like GloVe) truly capture semantic relationships, the "sentiment" of a word should be measurable by its proximity to known polar seeds in the embedding manifold.
Methodology: From Vectors to Polarity
The core innovation is a two-step scoring mechanism:
1. Positive-Word-Score (PWS)
The system defines two anchor points: a positive seed vector () and a negative seed vector (). The PWS of any word is defined as the difference between its similarity to these anchors:
This effectively projects a 100-dimensional vector onto a single scalar line representing "sentiment intensity."
2. Phrase-Sentiment-Score (PSS)
A sentence isn't just a bag of words; different parts of speech carry different "emotional loads." The authors introduce a weighted aggregation:
Figure 1: Visual mapping of word embeddings in the 2D plane showing semantic clusters of sentiment.
Experiments and Domain Adaptation
The model was stress-tested across three distinct domains: Sports, Politics, and Services.
Key Findings:
- Dimension Matters: Movement from 25D to 100D/200D embeddings generally yielded a 5-8% accuracy boost, as higher dimensions better resolve the nuance between neutral and slightly polar words.
- Corpus Scale: Using a 27B token corpus outperformed a 6B token corpus significantly, reinforcing the idea that "common sense" sentiment requires massive pre-training.
Table 1: Comparison of Accuracy across different embedding dimensions and vocabularies.
Critical Insight: The Power of Weighted POS
An interesting aspect of the research is the "Optimal Combination" of weights for Nouns, Verbs, Adverbs, and Adjectives. The researchers found that the best-performing models often assigned the highest weights to Verbs and Adjectives (weight = 4), while Nouns were often weighted lower (weight = 1). This aligns with linguistic intuition—actions and descriptions drive sentiment more than the entities themselves.
Conclusion & Limitations
The PWS-PSS framework offers a computationally efficient alternative to deep neural networks for sentiment tasks where labeled data might be scarce.
Limitations:
- The method relies heavily on the quality of the "Seed Vectors." If the seeds are biased, the entire scoring system shifts.
- Sarcasm remains a major hurdle, as the linear aggregation of PWS cannot easily capture the "negation effect" found in complex sentence structures.
Future Work: Integrating this vector-distance approach with attention-based mechanisms could allow for dynamic weighting of words based on their local context rather than static POS weights.
