CAN-GRU: Solving the "Yeah" Problem in Dialogue Emotion Recognition

CAN-GRU: a Hierarchical Model for Emotion Recognition in Dialogue

2020-10-01
Ting Jiang, Bing Xu, Tiejun Zhao, Sheng Li
Summary
Problem
Method
Results
Takeaways
Abstract

CAN-GRU is a hierarchical neural network designed for Emotion Recognition in Conversations (ERC). It combines a Convolutional self-Attention Network (CAN) for fine-grained utterance feature extraction with a GRU-based layer to capture inter-utterance dependencies, achieving SOTA-level performance on the Friends and EmotionPush datasets.

TL;DR

Emotion Recognition in Conversation (ERC) is notoriously difficult because emotion is context-dependent—a simple "Yeah" can signify joy, neutral agreement, or surprise. CAN-GRU addresses this by using a hierarchical architecture: a Convolutional self-Attention Network (CAN) to extract deep semantic features within a single sentence, and a Bidirectional GRU with Attention to track the emotional flow throughout the entire conversation.

The Motivation: Why Emotions are Hard to Track

Standard NLP models often treat sentences in isolation. However, in a dialogue, the emotional state is a moving target. The authors identify two primary pain points:

  1. N-gram Information vs. Context: Traditional CNNs are great at local patterns (n-grams) but miss long-range word relationships. Standard Attention is great at relationships but ignores the local structure.
  2. Long-term Dialogue Dependency: Conversations can be long; forgetting what was said three turns ago leads to incorrect classification of the current utterance's response.

The core insight of the paper is that we need to see both the local context within an utterance and the global context of the dialogue window.


Methodology: The Hierarchical Approach

Layer 1: The Convolutional self-Attention Network (CAN)

Instead of feeding raw word embeddings directly into an attention mechanism, CAN-GRU applies convolutions first. This "pre-processing" step ensures that the Query (Q), Key (K), and Value (V) matrices already contain local n-gram information.

Model Architecture

As shown in the architecture, the model computes two separate attention results ( and ) and performs an element-wise multiplication. This specific design allows the model to capture more complex interactions than a standard weighted average.

Layer 2: Dialogue Modeling with biGRUA

Once each utterance is compressed into a vector, the model passes it through a Bidirectional GRU.

  • The "Future" matters: In many cases, we only understand a person's current emotion after seeing their next reaction. Bi-directional processing allows the model to "look ahead."
  • Self-Attention on Top: To prevent the GRU from "forgetting" the beginning of a long dialogue, a self-attention layer is added to aggregate the hidden states () globally.

Experiments and Results: Setting New Baselines

The authors tested their model on two major benchmarks: Friends (TV show transcripts) and EmotionPush (social media logs).

Performance Leap

The results (using Unweighted Accuracy - UWA) show a clear trend: adding layers of sophistication consistently improves performance.

Experimental Results Comparison

Key Observations:

  • CAN-biGRUA achieved the highest scores (65.3% and 67.1%), significantly beating the 2018 EmotionX Challenge winner (CNN-DCNN).
  • The Power of BERT: When replacing GloVe with BERT embeddings, the performance skyrocketed to 84.1%. This proves that while the architecture is strong, the quality of the underlying word representations is still a massive bottleneck in ERC.

Case Study: Historical vs. Future Context

One of the most compelling parts of the paper is the visualization of attention weights.

Attention Weight Comparison

In the figure above, CAN-GRUA (unidirectional) failed to recognize "Anger" because it only looked at the past. However, CAN-biGRUA looked at the subsequent utterances (7 and 8), which were highly aggressive, allowing it to correctly infer the anger in the 6th utterance.


Technical Insights & Future Outlook

Takeaway 1: Convolutions as Feature Extractors for Attention Most modern Transformers skip convolutions. This paper argues that for short, punchy dialogue utterances, n-gram extraction via convolution before the attention dot-product significantly improves the "read" on local sentiment.

Takeaway 2: Limitations in Complex Scenarios Despite its strengths, the model still struggles when emotions shift rapidly or when sarcasm is involved. The authors noted that "sorry" can often be misclassified as "sadness" even when the speaker is being neutral or sarcastic, suggesting that deeper pragmatic understanding is still required.

Future Direction: The authors plan to integrate more "Emotion Shift" detection logic to handle conversations where the tone changes from joy to anger within a single turn.

Find Similar Papers

Try Our Examples

  • Search for recent papers on Emotion Recognition in Conversation (ERC) that utilize graph neural networks (GNNs) to model multi-party speaker dependencies beyond hierarchical GRUs.
  • Which paper first introduced the concept of Convolutional Self-Attention, and how does the CAN-GRU implementation differ in its QKV generation process?
  • Explore how hierarchical attention mechanisms from CAN-GRU have been applied to multi-modal sentiment analysis involving video and audio features.
Contents
CAN-GRU: Solving the "Yeah" Problem in Dialogue Emotion Recognition
1. TL;DR
2. The Motivation: Why Emotions are Hard to Track
3. Methodology: The Hierarchical Approach
3.1. Layer 1: The Convolutional self-Attention Network (CAN)
3.2. Layer 2: Dialogue Modeling with biGRUA
4. Experiments and Results: Setting New Baselines
4.1. Performance Leap
4.2. Case Study: Historical vs. Future Context
5. Technical Insights & Future Outlook