Harmonizing Words and Feelings: A Unified SSL Framework for Joint ASR-SER

Semi-Supervised Learning for Multimodal Speech and Emotion Recognition

2021-10-15
Yuanchao Li
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a PhD research plan for a joint Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) framework using multimodal features (acoustic, visual/lip-reading, and lexical). The core strategy involves a novel cross-task semi-supervised learning (SSL) approach to leverage massive amounts of unlabeled data, aiming to surpass current SOTA performance in emotion-aware interactive systems.

TL;DR

This research tackles the "data desert" in Speech Emotion Recognition (SER) by proposing a joint architecture with Automatic Speech Recognition (ASR). By utilizing multimodal features (audio, lip-reading, and lexical) and a hierarchical fusion strategy, the author aims to use large unlabeled datasets through a novel cross-task Semi-Supervised Learning (SSL) approach. Early results show a significant jump in accuracy (up to ~66%) when using hierarchical attention over standard methods.

The Problem: Small Data, Big Emotions

In the world of AI, ASR is a mature giant, trained on thousands of hours of data like LibriSpeech. SER, however, is a niche orphan. The gold-standard IEMOCAP dataset contains a measly 12 hours of labeled audio. This creates two major issues:

  1. Lack of Robustness: Models overfit to small sets and fail in the "wild."
  2. The Multimodal Gap: While we know that what we say (lexical) and how we look (visual) matters as much as how we sound (acoustic), fusing these effectively without massive labels is historically difficult.

The Insight: Cross-Task Synergy

The author's PhD plan rests on a brilliant intuition: ASR and SER should not be silos.

  • Lexical Bridge: ASR provides the "what," which is essential context for "how" (emotion).
  • Visual Stabilizer: Lip-reading doesn't change based on emotion as much as pitch does, making it a "ground truth" anchor to help ASR remain accurate even when a speaker is shouting or crying.
  • Mutual Supervision: In an SSL setting, if the ASR is confident about the text and the SER is confident about the emotion, they can provide high-quality "pseudo-labels" for unlabeled data, creating a virtuous cycle of learning.

Methodology: Hierarchical & Attentional Fusion

The framework moves away from "shallow fusion" (simply concatenating vectors). Instead, it mirrors human auditory processing—moving from low-level frames to high-level semantics.

Model Architecture Concept

Core Components:

  • Acoustic: Uses wav2vec 2.0 for self-supervised utterance-level representations.
  • Visual: Employs spatio-temporal ResNets for character-level lip-reading.
  • Lexical: Extracts hidden states from ASR instead of raw text to maintain robustness against transcription errors.
  • Fusion: Uses Co-Attention mechanisms to interchange "Key-Value" pairs between modalities, ensuring the model attends to the most emotionally salient parts of the speech.

Experimental Validation

The preliminary results provide strong evidence for the "Hierarchical" hypothesis. By processing lower-level features through CNN-BLSTM layers before fusing them with high-level BERT embeddings, the model achieves superior performance.

Feature TypeFusion ApproachWeighted Accuracy
wav2vec (Audio only)None57.67%
BERT (Text only)None50.75%
wav2vec + BERTHierarchical Shallow Fusion66.63%
wav2vec + BERTHierarchical Attentional Fusion62.06%

(Note: While HSF performed slightly better than HAF in this snapshot, the author notes that attentional mechanisms are more flexible for the planned tri-modal expansion.)

Performance Comparison Table

Critical Analysis & Future Outlook

The Good: The transition to self-supervised backbones (wav2vec 2.0) is a major step forward. The emphasis on "uncertainty modeling" in SSL is the right way to handle the noise prevalent in emotional datasets.

The Challenge: Joint training of ASR and SER is notoriously difficult because emotion often degrades ASR performance. The author’s plan to use lip-reading to "bridge" this gap is clever but computationally expensive.

Takeaway: This work signals a shift in affect computing. We are moving away from building "emotion classifiers" and toward building "perceptive agents" that understand language and feelings as a unified signal.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2022-2024 that implement joint ASR and SER training using wav2vec 2.0 or Whisper backbones.
  • Which study first introduced the concept of 'cross-task semi-supervised learning' for speech tasks, and how does this paper's uncertainty modeling differ?
  • Look for research applying lip-reading (visual-only speech recognition) to enhance emotional sentiment analysis in noisy or multi-speaker environments.
Contents
Harmonizing Words and Feelings: A Unified SSL Framework for Joint ASR-SER
1. TL;DR
2. The Problem: Small Data, Big Emotions
3. The Insight: Cross-Task Synergy
4. Methodology: Hierarchical & Attentional Fusion
4.1. Core Components:
5. Experimental Validation
6. Critical Analysis & Future Outlook