Ensemble Temporal Learning: A Multi-Model Approach to Emotion Recognition

Multiple Models Using Temporal Feature Learning for Emotion Recognition

2020-09-17
Hoang Manh Hung, Soo-Hyung Kim, Hyung-Jeong Yang, Guee-Sang Lee
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a multi-model ensemble framework for video-based emotion recognition using facial expression features. By combining LSTM with Self-Attention, WaveNet, and SVM via late fusion, the method achieved state-of-the-art performance on the MuSe-CaR dataset, particularly outperforming baselines in arousal prediction.

Executive Summary

TL;DR: This research introduces a hybrid framework that leverages the strengths of multiple architectures—LSTM, WaveNet, and SVM—to tackle the complexity of facial emotion recognition in video sequences. By focusing on temporal feature learning, the authors achieved significant gains in predicting emotional intensity levels.

Positioning: This work is a "SOTA-competitive" refinement that emphasizes the power of model ensembles. It proves that even without multi-modal inputs (like audio or text), a sophisticated visual temporal pipeline can outperform broader multi-modal baselines in specific dimensions like physiological arousal.

Motivation: The Complexity of "In-the-Wild" Emotions

Recognizing emotions from real-world video (such as car reviews in the MuSe-CaR dataset) is notoriously difficult. Unlike static images, video requires a model to understand how a smile fades or how a brow furrows over time.

The authors identified a gap in existing methods: single-stream models often overfit to specific spatial features or fail to grasp the hierarchical nature of temporal changes. Their insight was to treat the temporal sequence as a signal that can be parsed through different "lenses"—the sequential memory of an LSTM, the multi-scale receptive fields of a WaveNet, and the statistical boundaries of an SVM.

Methodology: The Triple-Threat Architecture

The proposed pipeline follows a sophisticated multi-stage process:

1. Robust Feature Extraction

Before modeling temporal dynamics, the system ensures high-quality spatial inputs:

  • MTCNN: Used for face detection and landmark localization.
  • Face Alignment: Normalizes faces to a standard (112, 112) resolution to reduce noise from head pose variations.
  • ResNet50: Pre-trained on MS-Celeb-1M, this serves as the "backbone" to generate 512-dimensional embedding vectors for every frame.

The pre-processing for extracting the feature vectors.

2. Multi-Model Temporal Learning

The core innovation lies in parallel processing:

  • LSTM + Self-Attention: Focuses on long-term global context, ensuring the model remembers the "emotional baseline" of the subject.
  • WaveNet: Utilizes stacked dilated convolutions. This is a clever adaptation from audio processing to vision, allowing the model to have a massive receptive field to catch subtle, fast-moving facial micro-expressions.
  • SVM: Acts as a "statistical anchor," using ANOVA F-value feature selection to focus on the most discriminative dimensions of the ResNet embeddings.

Our proposed system combined 3 models.

Experimental Analysis

The framework was tested on the MuSe-CaR dataset, which consists of 37 hours of video reviews. The evaluation focused on Valence (positivity/negativity) and Arousal (intensity).

Performance Gains

The results show a clear advantage in capturing "Arousal":

MethodFeaturesValenceArousalCombined
Baseline (MMT)Text+Audio+Video37.6546.5842.12
ProposedVideo Only (VG)38.1150.1744.14

Experimental Results Comparison

The 4% jump in Arousal is particularly impressive because the baseline used multiple modalities (including audio features which are typically strong indicators of arousal), while the proposed method relied solely on visual facial features.

Critical Insights & Future Outlook

Why did it work? The success of WaveNet in this context suggests that dilated convolutions are highly effective at capturing the "rhythm" of facial movements. By combining this with LSTM’s attention-based memory, the model effectively covers both the "what" (spatial) and the "when" (temporal) of human emotion.

Limitations: The paper utilizes late fusion, which simply combines the final decisions. A "feature-level" early fusion or a cross-attention mechanism between the three models might have yielded even more nuanced representations.

Conclusion: This research reinforces the principle that in complex affective computing tasks, an ensemble of specialized temporal learners is often superior to a single monolithic network. For practitioners, the takeaway is clear: don't just build deeper; build more diversely.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize WaveNet's dilated convolutions for video-based temporal feature extraction outside of audio synthesis.
  • Which study first introduced the 'late fusion' strategy for combining deep learning and traditional machine learning models in affective computing?
  • How do current SOTA methods on the MuSe-CaR dataset integrate facial expressions with audio and text modalities to improve valence prediction?
Contents
Ensemble Temporal Learning: A Multi-Model Approach to Emotion Recognition
1. Executive Summary
2. Motivation: The Complexity of "In-the-Wild" Emotions
3. Methodology: The Triple-Threat Architecture
3.1. 1. Robust Feature Extraction
3.2. 2. Multi-Model Temporal Learning
4. Experimental Analysis
4.1. Performance Gains
5. Critical Insights & Future Outlook