Enhancing Conversational Emotion Recognition: The Power of Weighted Historical Context
13139_Emotional State Estimation by Dialogue History and Sentence Distributed Representation.
This paper introduces a novel conversational emotion recognition framework presented at CCIS 2019, which utilizes historical utterance vectors ( past utterances) to augment the current feature space. By implementing various weighting strategies—specifically "mul_log" (logarithmic decay)—the method achieves a significant performance boost in multi-speaker dialogue scenarios.
TL;DR
In the context of multi-turn dialogues, emotions are rarely isolated events; they are influenced by what was said moments ago. This paper presents a method to enhance the representation of a current utterance by incorporating a weighted history of previous utterances. By applying a logarithmic decay to past information, the proposed "mul_log" method significantly outperforms standard RNNs and LSTMs in predicting emotional states across diverse conversational scenarios.
The Challenge: Contextual Emotional Inertia
Most early emotion recognition systems struggled with the "context gap." Recognizing internal states in a vacuum is difficult because human dialogue is characterized by Emotional Inertia—a phenomenon where a speaker's current feeling is a mix of their baseline state and the "ripples" left by previous interactions. Existing RNN-based approaches often suffer from vanishing gradients or fail to explicitly model how the relevance of past utterances diminishes over time.
Methodology: Beyond Simple Recurrence
The authors propose a series of mathematical formulations to augment the current utterance vector . The core idea is to transform into by adding a localized context vector.
Key Mathematical Formulations
The paper explores several ways to weight the influence of the past ( steps back):
- Inverse Linear (mul): Weights the past by .
- Logarithmic Decay (mul_log):
Insight: Use logarithmic scaling to model a more natural "fade" of memory. - Speaker-Aware Modulation (other):
Insight: Differentiation between the speaker's own past emotions and those of their interlocutors.
System Architecture
The architecture (shown below) illustrates how the history window is processed to enrich the feature set before final classification.

Experimental Analysis
The researchers tested their methods on 10 distinct scenarios (SC-01 to SC-10) involving varying numbers of speakers and utterances.
SOTA Comparison
As shown in the results table, the "hard-coded" weighting mechanisms (specifically mul and mul_log) surprisingly outperformed learnable recurrent architectures like LSTM and GRU.
| Method | M=5 | M=10 | M=15 |
|---|---|---|---|
| mul_log | 54.7% | 65.4% | 65.9% |
| LSTM | 53.6% | 52.9% | 55.5% |
| GRU | 52.6% | 54.4% | 53.9% |
| None (Baseline) | - | 53.4% | - |

Emotion-Specific Performance
The system showed a high precision for "Neutral" (78.6%) and "Surprise" (76.6%), whereas "Positive" emotions proved more challenging to distinguish (55.7%), likely due to the linguistic overlap between positive sentiment and neutral politeness in professional scenarios.
Critical Insight & Conclusion
This work demonstrates that explicit temporal modeling is often more critical than model depth in short-to-medium dialogue contexts. While the industry has shifted toward Transformers, the finding that logarithmic decay effectively simulates human emotional memory remains a valuable heuristic for building efficient, low-latency conversational agents.
Future Outlook: Integrating these decay-based features as a "Temporal Bias" in Transformer attention heads could potentially combine the best of both worlds—global context and natural emotional decay.
