Enhancing Conversational Emotion Recognition: The Power of Weighted Historical Context

13139_Emotional State Estimation by Dialogue History and Sentence Distributed Representation.

Summary
Problem
Method
Results
Takeaways

This paper introduces a novel conversational emotion recognition framework presented at CCIS 2019, which utilizes historical utterance vectors ( past utterances) to augment the current feature space. By implementing various weighting strategies—specifically "mul_log" (logarithmic decay)—the method achieves a significant performance boost in multi-speaker dialogue scenarios.

TL;DR

In the context of multi-turn dialogues, emotions are rarely isolated events; they are influenced by what was said moments ago. This paper presents a method to enhance the representation of a current utterance by incorporating a weighted history of previous utterances. By applying a logarithmic decay to past information, the proposed "mul_log" method significantly outperforms standard RNNs and LSTMs in predicting emotional states across diverse conversational scenarios.

The Challenge: Contextual Emotional Inertia

Most early emotion recognition systems struggled with the "context gap." Recognizing internal states in a vacuum is difficult because human dialogue is characterized by Emotional Inertia—a phenomenon where a speaker's current feeling is a mix of their baseline state and the "ripples" left by previous interactions. Existing RNN-based approaches often suffer from vanishing gradients or fail to explicitly model how the relevance of past utterances diminishes over time.

Methodology: Beyond Simple Recurrence

The authors propose a series of mathematical formulations to augment the current utterance vector . The core idea is to transform into by adding a localized context vector.

Key Mathematical Formulations

The paper explores several ways to weight the influence of the past ( steps back):

  1. Inverse Linear (mul): Weights the past by .
  2. Logarithmic Decay (mul_log): Formula Insight: Use logarithmic scaling to model a more natural "fade" of memory.
  3. Speaker-Aware Modulation (other): Formula Insight: Differentiation between the speaker's own past emotions and those of their interlocutors.

System Architecture

The architecture (shown below) illustrates how the history window is processed to enrich the feature set before final classification.

Model Architecture Placeholder

Experimental Analysis

The researchers tested their methods on 10 distinct scenarios (SC-01 to SC-10) involving varying numbers of speakers and utterances.

SOTA Comparison

As shown in the results table, the "hard-coded" weighting mechanisms (specifically mul and mul_log) surprisingly outperformed learnable recurrent architectures like LSTM and GRU.

MethodM=5M=10M=15
mul_log54.7%65.4%65.9%
LSTM53.6%52.9%55.5%
GRU52.6%54.4%53.9%
None (Baseline)-53.4%-

Performance Data

Emotion-Specific Performance

The system showed a high precision for "Neutral" (78.6%) and "Surprise" (76.6%), whereas "Positive" emotions proved more challenging to distinguish (55.7%), likely due to the linguistic overlap between positive sentiment and neutral politeness in professional scenarios.

Critical Insight & Conclusion

This work demonstrates that explicit temporal modeling is often more critical than model depth in short-to-medium dialogue contexts. While the industry has shifted toward Transformers, the finding that logarithmic decay effectively simulates human emotional memory remains a valuable heuristic for building efficient, low-latency conversational agents.

Future Outlook: Integrating these decay-based features as a "Temporal Bias" in Transformer attention heads could potentially combine the best of both worlds—global context and natural emotional decay.

Find Similar Papers

Try Our Examples

  • Search for recent papers that compare explicit context weighting (like logarithmic decay) against Transformer-based attention mechanisms in conversational emotion recognition (ERC).
  • Which study first introduced the concept of "Emotional Inertia" in dialogue systems, and how does this paper's vector-based approach align with that theory?
  • How can weighted historical emotion vectors be integrated into multi-modal systems involving both audio and text features for real-time sentiment analysis?
Contents
Enhancing Conversational Emotion Recognition: The Power of Weighted Historical Context
1. TL;DR
2. The Challenge: Contextual Emotional Inertia
3. Methodology: Beyond Simple Recurrence
3.1. Key Mathematical Formulations
3.2. System Architecture
4. Experimental Analysis
4.1. SOTA Comparison
4.2. Emotion-Specific Performance
5. Critical Insight & Conclusion