LSSED: Bridging the Gap Between Lab-Based and Real-World Speech Emotion Recognition

LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition

2021-05-13
Weiquan Fan, Xiangmin Xu, Xiaofen Xing, Weidong Chen, Dongyan Huang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces LSSED (Large-Scale Speech Emotion Dataset), a massive English speech emotion corpus containing over 200 hours of naturalistic data from 820 subjects. The authors also propose PyResNet, a pyramid-convolution-based architecture, setting a new SOTA benchmark for large-scale speech emotion recognition (SER).

TL;DR

Speech Emotion Recognition (SER) has long been stifled by "small-data" overfitting. This paper breaks that bottleneck by introducing LSSED, a large-scale dataset featuring over 200 hours of natural English speech from 820 subjects. Alongside this, the authors propose PyResNet, a multi-scale architecture that outperforms standard ResNets and demonstrates that emotional pre-training is far more effective than traditional ASR pre-training for downstream tasks like depression detection.

Problem & Motivation: The Overfitting Trap

In the world of SER, we have reached a plateau. While many models boast high accuracy on datasets like IEMOCAP or EMODB, they frequently fail when deployed in the wild. The reason is twofold:

  1. Data Scarcity: Previous "standard" datasets are tiny. IEMOCAP has only 10 speakers and ~13 hours of audio. Models simply memorize these specific voices instead of learning the "universal language" of emotion.
  2. Semantic Bias: Most pre-trained speech models come from Automatic Speech Recognition (ASR). ASR models are designed to ignore "noise" (which includes emotional prosody) to focus strictly on text. For mental health analysis (e.g., detecting depression), we need the exact opposite: we need the "how" it was said, not just the "what."

Methodology: Scalability and Multi-Scale Intuition

The authors address these gaps with a two-pronged approach: the LSSED Dataset and the PyResNet Architecture.

1. The LSSED Dataset

Unlike acted datasets, LSSED captures spontaneous utterances across 11 labels (Anger, Happiness, Sadness, etc.). It reaches 206 hours, making it roughly 15-20 times larger than its predecessors.

2. PyResNet: Capturing the "Vibe"

To process this data, the authors developed PyResNet. The core technical innovation is the integration of Pyramid Convolutions.

  • Physical Intuition: Emotional cues in speech (like a sudden break in tone or a long-term depressive drone) happen at different time scales. Standard ResNets utilize fixed-size kernels. Pyramid convolutions use multiple kernel sizes simultaneously, allowing the model to capture both micro-prosody and macro-intonation.
  • Frequency Preservation: By replacing Global Average Pooling (GAP) with pooling only in the time dimension, the model retains the frequency resolution necessary to distinguish subtle acoustic shifts.

Model Performance Comparison Table: LSSED results showing PyResNet outperforming standard backbones and prior SOTA algorithms.

Experiments: Real-World Generalization

The authors conducted a "stress test" by training on one dataset and testing on another.

  • Small Large: A model trained on IEMOCAP lost nearly 60% accuracy when tested on LSSED.
  • Large Small: A model trained on LSSED only lost 11.9% accuracy when tested on IEMOCAP.

This proves that LSSED contains a much more robust representation of human emotion that "covers" the distribution of smaller, artificial datasets.

Downstream Success: Depression Detection

In a transfer learning experiment on the DAIC-WOZ depression database, PyResNet (pre-trained on LSSED) achieved a WA of 0.714, significantly outperforming the industry-standard ESPNet (pre-trained on ASR) at 0.657. This validates the hypothesis that acoustic-emotional pre-training is superior to linguistic pre-training for medical and psychological applications.

Confusion Matrices Figure: PyResNet shows improved prediction density for primary emotions, though "Neutral" bias remains a challenge.

Critical Analysis & Conclusion

Takeaway: LSSED is a significant milestone for SER, providing the community with a "large-scale" foundation similar to what ImageNet did for CV or BERT did for NLP.

Limitations: Despite the scale, the dataset reflects a common "real-world" problem: Class Imbalance. Neutral samples dominate the distribution. As shown in the confusion matrices, even the best models still struggle with a "Neutral bias," often misclassifying subtle emotions as neutral.

Future Outlook: The release of pre-trained models on LSSED opens the door for high-precision mental health monitoring tools that can work on raw audio without needing privacy-invasive text transcripts. The next step for the field will likely be combining this large-scale supervised data with self-supervised (SSL) frameworks to further push the boundaries of acoustic feature extraction.

Find Similar Papers

Try Our Examples

  • Find recent balance-sampling or cost-sensitive learning papers addressing the class imbalance between 'Neutral' and rare emotions like 'Fear' or 'Surprise' in spontaneous speech datasets.
  • Which paper first introduced the concept of Pyramid Convolutions in the context of computer vision, and how does this paper adapt that inductive bias for 1D/2D audio spectrograms?
  • Search for recent studies that utilize self-supervised learning (SSL) backbones like Hubert or Wav2Vec2.0 on the LSSED dataset to see if they outperform the supervised PyResNet approach.
Contents
LSSED: Bridging the Gap Between Lab-Based and Real-World Speech Emotion Recognition
1. TL;DR
2. Problem & Motivation: The Overfitting Trap
3. Methodology: Scalability and Multi-Scale Intuition
3.1. 1. The LSSED Dataset
3.2. 2. PyResNet: Capturing the "Vibe"
4. Experiments: Real-World Generalization
4.1. Downstream Success: Depression Detection
5. Critical Analysis & Conclusion