CAN-GRU: Solving the "Yeah" Problem in Dialogue Emotion Recognition
CAN-GRU: a Hierarchical Model for Emotion Recognition in Dialogue
CAN-GRU is a hierarchical neural network designed for Emotion Recognition in Conversations (ERC). It combines a Convolutional self-Attention Network (CAN) for fine-grained utterance feature extraction with a GRU-based layer to capture inter-utterance dependencies, achieving SOTA-level performance on the Friends and EmotionPush datasets.
TL;DR
Emotion Recognition in Conversation (ERC) is notoriously difficult because emotion is context-dependent—a simple "Yeah" can signify joy, neutral agreement, or surprise. CAN-GRU addresses this by using a hierarchical architecture: a Convolutional self-Attention Network (CAN) to extract deep semantic features within a single sentence, and a Bidirectional GRU with Attention to track the emotional flow throughout the entire conversation.
The Motivation: Why Emotions are Hard to Track
Standard NLP models often treat sentences in isolation. However, in a dialogue, the emotional state is a moving target. The authors identify two primary pain points:
- N-gram Information vs. Context: Traditional CNNs are great at local patterns (n-grams) but miss long-range word relationships. Standard Attention is great at relationships but ignores the local structure.
- Long-term Dialogue Dependency: Conversations can be long; forgetting what was said three turns ago leads to incorrect classification of the current utterance's response.
The core insight of the paper is that we need to see both the local context within an utterance and the global context of the dialogue window.
Methodology: The Hierarchical Approach
Layer 1: The Convolutional self-Attention Network (CAN)
Instead of feeding raw word embeddings directly into an attention mechanism, CAN-GRU applies convolutions first. This "pre-processing" step ensures that the Query (Q), Key (K), and Value (V) matrices already contain local n-gram information.

As shown in the architecture, the model computes two separate attention results ( and ) and performs an element-wise multiplication. This specific design allows the model to capture more complex interactions than a standard weighted average.
Layer 2: Dialogue Modeling with biGRUA
Once each utterance is compressed into a vector, the model passes it through a Bidirectional GRU.
- The "Future" matters: In many cases, we only understand a person's current emotion after seeing their next reaction. Bi-directional processing allows the model to "look ahead."
- Self-Attention on Top: To prevent the GRU from "forgetting" the beginning of a long dialogue, a self-attention layer is added to aggregate the hidden states () globally.
Experiments and Results: Setting New Baselines
The authors tested their model on two major benchmarks: Friends (TV show transcripts) and EmotionPush (social media logs).
Performance Leap
The results (using Unweighted Accuracy - UWA) show a clear trend: adding layers of sophistication consistently improves performance.

Key Observations:
- CAN-biGRUA achieved the highest scores (65.3% and 67.1%), significantly beating the 2018 EmotionX Challenge winner (CNN-DCNN).
- The Power of BERT: When replacing GloVe with BERT embeddings, the performance skyrocketed to 84.1%. This proves that while the architecture is strong, the quality of the underlying word representations is still a massive bottleneck in ERC.
Case Study: Historical vs. Future Context
One of the most compelling parts of the paper is the visualization of attention weights.

In the figure above, CAN-GRUA (unidirectional) failed to recognize "Anger" because it only looked at the past. However, CAN-biGRUA looked at the subsequent utterances (7 and 8), which were highly aggressive, allowing it to correctly infer the anger in the 6th utterance.
Technical Insights & Future Outlook
Takeaway 1: Convolutions as Feature Extractors for Attention Most modern Transformers skip convolutions. This paper argues that for short, punchy dialogue utterances, n-gram extraction via convolution before the attention dot-product significantly improves the "read" on local sentiment.
Takeaway 2: Limitations in Complex Scenarios Despite its strengths, the model still struggles when emotions shift rapidly or when sarcasm is involved. The authors noted that "sorry" can often be misclassified as "sadness" even when the speaker is being neutral or sarcastic, suggesting that deeper pragmatic understanding is still required.
Future Direction: The authors plan to integrate more "Emotion Shift" detection logic to handle conversations where the tone changes from joy to anger within a single turn.
