AVEC 2018: Pushing the Boundaries of Mental Health and Cross-Cultural Affect Recognition

AVEC 2018 Workshop and Challenge: Bipolar Disorder and Cross-Cultural Affect Recognition

2018-10-15
Fabien Ringeval, Björn Schuller, Michel Valstar, Roddy Cowie, Heysem Kaya, Maximilian Schmitt, Shahin Amiriparian, Nicholas Cummins, Denis Lalanne, Adrien Michaud, Elvan Ciftçi, Hüseyin Güleç, Albert Ali Salah, Maja Pantic, M. Pantic
Summary
Problem
Method
Results
Takeaways
Abstract

The AVEC 2018 workshop introduces three standardized sub-challenges for automatic audiovisual analysis: Bipolar Disorder (BDS), Cross-cultural Emotion (CES), and Gold-standard Emotion generation (GES). It establishes a robust benchmark using the Concordance Correlation Coefficient (CCC) and Unweighted Average Recall (UAR) to evaluate multimodal machine learning methods.

TL;DR

The 8th Audio/Visual Emotion Challenge (AVEC 2018) marks a significant shift in affective computing by introducing the first-ever competition for Bipolar Disorder classification. Beyond clinical health, it tackles the "in-the-wild" challenge of cross-cultural emotion recognition (German vs. Hungarian) and the theoretical problem of generating a unified "gold-standard" from subjective human ratings.

Background: Why AVEC 2018 Matters

Affective computing has long struggled with "lab-grown" data that doesn't survive contact with reality. AVEC 2018 bridges this gap by focusing on:

  1. Clinical Utility: Moving from general depression to the specific manic/hypomanic phases of Bipolar Disorder.
  2. Cultural Robustness: Testing if an AI trained on one culture (German) can understand another (Hungarian).
  3. Label Integrity: Developing math-driven ways to reconcile the fact that different human annotators rarely agree on the exact timing or intensity of an emotion.

Methodology: The Core Architecture

The organizers provided a multi-layered baseline architecture incorporating supervised, semi-supervised, and unsupervised feature extraction.

1. The Deep Spectrum Approach (Unsupervised)

A standout novelty is the use of Deep Spectrum features. Instead of manually defining "pitch" or "energy," speech is converted into mel-spectrogram images. These images are fed into a pre-trained AlexNet (originally for image recognition), using the activations from the fc7 layer as feature vectors. This treats audio as a visual pattern-recognition problem.

2. Time-Dependent Modeling

For emotion recognition, the baseline uses a 2-layer LSTM-RNN. This is critical because emotions are not static; the "valence" (positivity/negativity) of a sentence depends on the temporal context of what was said seconds before.

Model Architecture and Task Overview

Experiments and Results

The results highlight the difficulty of "in-the-wild" cross-cultural tasks compared to controlled environments.

  • Bipolar Disorder (BDS): The unsupervised Deep Spectrum features (58.20% UAR) outperformed traditional expert-knowledge features (eGeMAPS) on the development set, suggesting that deep learning can find "hidden" biomarkers of mania that human-defined features miss.
  • Cross-Cultural (CES): While Arousal and Valence yielded reasonable Concordance Correlation Coefficients (CCC), "Liking" was poorly predicted. This suggests that "liking" something is more culture-specific or linguistically driven than the universal physical signals of "excitement" (arousal).

Performance Summary Table

Below is a snippet of the baseline performance across modalities:

Baseline Results Comparison

Critical Insights & Future Outlook

The "Reaction Time" Problem: The GES (Gold-standard) task highlights a persistent issue: humans have a physical delay (reaction time) when annotating video. The AVEC 2018 baseline compensates for this by shifting labels in time to maximize correlation—a "calibration" step that is now essential for any time-continuous affect model.

Limitations: The primary limitation noted is the omission of linguistic transcriptions in the baseline. In real-world scenarios, what is said is often as important as how it is said.

Takeaway for Researchers

If you are building AI for healthcare or social robotics, the lesson of AVEC 2018 is clear: unsupervised feature learning (like Deep Spectrum) is closing the gap with expert-tuned features, and multimodal fusion (combining video and audio) is no longer optional—it is the prerequisite for dealing with human diversity and cultural nuances.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Spectrum or CNN-based spectrogram features for mental health monitoring beyond Bipolar Disorder.
  • Which original studies established the Young Mania Rating Scale (YMRS) as a ground truth for clinical audiovisual computing tasks?
  • Find research that applies the Evaluation Weighted Estimator (EWE) or other advanced label fusion techniques for time-continuous emotion annotation in different languages.
Contents
AVEC 2018: Pushing the Boundaries of Mental Health and Cross-Cultural Affect Recognition
1. TL;DR
2. Background: Why AVEC 2018 Matters
3. Methodology: The Core Architecture
3.1. 1. The Deep Spectrum Approach (Unsupervised)
3.2. 2. Time-Dependent Modeling
4. Experiments and Results
4.1. Performance Summary Table
5. Critical Insights & Future Outlook
5.1. Takeaway for Researchers