Fusion of Lexicons and Statistics: Mastering Weibo Emotion Classification
Emotion Classification of Chinese Microblog Text via Fusion of BoW and eVector Feature Representations
This paper presents a hybrid emotion classification system for Chinese Sina Weibo texts, categorizing content into seven basic emotions (anger, disgust, fear, happiness, like, sadness, surprise). The authors propose a novel "eVector" feature representation based on a custom-built emotion lexicon and fuse it with a traditional Bag-of-Words (BoW) SVM baseline.
TL;DR
Researchers from Renmin University of China developed a specialized system for "fine-grained" sentiment analysis on Sina Weibo. By combining a traditional Bag-of-Words (BoW) approach with a novel eVector (Emotion Vector) representation—derived from a uniquely weighted emotion lexicon—the system achieves superior performance in identifying seven distinct emotional states in short, informal Chinese text.
Background & Motivation: Why Weibo is a Hard Nut to Crack
Traditional sentiment analysis often focuses on "polarity" (positive vs. negative). However, understanding why a user is upset (is it anger, disgust, or fear?) requires deeper nuance. Weibo presents four specific hurdles:
- Limited Length: 140-character limits offer sparse data.
- Linguistic Complexity: Chinese sentence structures and web slang (e.g., “跪了” - kneeling) change meanings frequently.
- Confusable Emotions: "Like" and "Happiness" are often entwined.
- Symbolic Language: High reliance on emoticons like "[抓狂]" (maddened) and punctuation.
Methodology: The Power of the eVector
The authors suggest that while BoW is a strong statistical baseline, it misses the "emotional weight" inherent in specific words. They introduce a two-system fusion:
1. The Baseline (BoW + SVM)
Using the Jieba segmenter, the authors extract not just adjectives, but nouns and verbs. Unique to this baseline is the inclusion of a specialized vocabulary for emotion expressions and punctuation, which is weighted more heavily than standard text.
2. The Contrast System (eVector)
The researchers built an emotion lexicon by categorizing words into three types: Emotional, Common, and Uncommon Neutral. They used a specific formula to rank words, ensuring that words occurring frequently in one emotion but rarely in others received the highest weight.

This resulted in a 7-dimensional vector representing the text, where each dimension corresponds to one of the seven target emotions.
Experiments and Results
The study evaluated the systems on both Document Level and Sentence Level.
Key Findings:
- Emoticons are King: As shown in the ablation study (Table 11), a model using only words (weight 1.0/0.0) performed significantly worse than a model incorporating expressions and punctuation.
- Complementary Gains: The fusion of BoW and eVector consistently outperformed either system alone, proving that statistical word distributions and lexicon-based emotional mapping capture different "signals" in the noise of social media.

Confusion Matrix Insights
The experiment revealed that "None" (neutral text) is the most difficult class to distinguish, often confused with "Like." Similarly, "Anger" and "Disgust" frequently overlap, suggesting that future models might benefit from hierarchical classification or label smoothing.
Critical Analysis & Conclusion
This work highlights a critical transition in NLP from pure statistics to feature fusion. While the eVector approach is simpler than today's Large Language Models (LLMs), its logic—weighting words based on their discriminative power across emotions—remains a core principle in affective computing.
Takeaway for Practitioners: When dealing with informal Chinese text, don't just look at the characters. The "meta-language" (punctuation, icons, and slang) often carries the bulk of the emotional payload.
Future Outlook: The authors suggest that improving the initial "emotion detection" (sentimental vs. non-sentimental) is the next frontier, as it remains the bottleneck for overall accuracy.
