Decoding the "Why" Behind the "What": Rule-Based Emotion Cause Detection in Chinese Micro-blogs
A rule-based approach to emotion cause detection for Chinese micro-blogs
This paper presents a rule-based framework for Emotion Cause Detection (ECD) specifically tailored for Chinese micro-blogs. It introduces the ECOCC (Emotion-Cause-OCC) model to categorize 22 fine-grained emotions and utilizes a Bayesian probability approach integrated with multi-linguistic features to identify specific cause components and their proportions.
TL;DR
While sentiment analysis identifies what someone feels, Emotion Cause Detection (ECD) seeks the underlying trigger. This paper introduces the ECOCC model, a rule-based framework for Chinese micro-blogs that maps 22 fine-grained emotions to their causal events. By leveraging Bayesian probability and multi-language features (emoticons, negations, etc.), the authors significantly outperform existing linguistic-cue baselines, reaching a precision of over 82% in component analysis.
Context & Motivation: Moving Beyond Polarities
Most sentiment analysis tools categorize text into simple buckets: Positive, Negative, or Neutral. However, for a brand manager or a government entity, knowing a crowd is "Angry" isn't enough—they need to know if the anger is directed at a product defect, a specific event, or an agent's action.
The challenge in Chinese micro-blogs (like Sina Weibo) lies in the brevity and complexity of the language. Users often omit subjects, use heavy sarcasm, or rely on emoticons to flip the meaning of a sentence entirely. Existing methods relied on rigid linguistic cues that failed to capture the psychological cognitive process of emotion triggering.
Methodology: The ECOCC Framework
The researchers developed the ECOCC (Emotion-Cause-OCC) model, an evolution of the classical OCC cognitive psychology model. It categorizes 22 fine-grained emotions into three primary branches:
- Results of Events: (e.g., Hope, Joy, Distress, Fear)
- Actions of Agents: (e.g., Pride, Shame, Admiration, Reproach)
- Aspects of Objects: (e.g., Liking, Disliking)
1. Rule-Based Extraction
The system identifies "internal events" (direct triggers) and "external events" (hashtags like #Topic#). Using Dependency Parsing and Semantic Role Labeling (SRL), it decomposes sentences into triples: .

2. Bayesian Intensity Scoring
To determine which cause is most significant when multiple triggers exist, the authors use Bayesian Probability. They calculate an "Emotion Intensity Score" () influenced by:
- Emoticons: Mapped to intensity through a co-occurrence graph.
- Degreed Adverbs: Intensifying or weakening the emotion via exponential functions ().
- Negations: Handling complex double negations where .
- Conjunctions: Determining if the focus (and thus the cause) lies before or after a "but" (但是) or "because" (å› ä¸º).

Experimental Battleground
The model was tested on a dataset of 16,371 Sina Weibo posts. The authors benchmarked their rule-based approach against two prominent baselines (Lee et al. and Li & Xu).
Performance Gains
- Baseline F-score: 69.99% (using only keywords)
- Full Feature F-score (EW + ALL): 75.46%
- Extraction Accuracy: 65.51%, representing a 12.95% improvement over traditional linguistic-cue methods.

The data reveals that emoticons and negation words provide the highest uplift in accuracy. This aligns with human intuition: in short-form social media, a single "😡" or a "not" carries more weight than the actual nouns in the sentence.
Critical Insight: Why Rules Still Matter
In the age of Large Language Models (LLMs), one might ask: why use a rule-based system? This paper reminds us that interpretability and linguistic structure are paramount in psychological mining. By using a Bayesian approach tied to formal cognitive categories, we don't just get a prediction—we get a structured explanation of the human psyche behind the keyboard.
Future Outlook
While the recall currently suffers from the extreme brevity of micro-blogs (where some posts contain only an emoji), future iterations could integrate Temporal Analysis—tracking how a cause event evolves over time. This has massive implications for Public Emergency Management and Precision Marketing, allowing systems to intervene or recommend products based on the specific origin of a user's sentiment.
Takeaway: To truly understand online public opinion, we must look past the "sentiment" and model the "triggering event" as a structural relationship between agents, actions, and objects.
