PCASS: Aligning Latent Features for Cross-Corpus Speech Emotion Recognition
Unsupervised domain adaptation for speech emotion recognition using PCANet
The paper introduces PCASS, an unsupervised domain adaptation framework for Speech Emotion Recognition (SER) based on the PCANet architecture. By extracting and aligning domain-shared and domain-specific latent features through an interpolating path, it achieves SOTA results on the FAU Aibo Emotion Corpus (FAU AEC).
TL;DR
Recognizing emotions across different datasets (cross-corpus) is notoriously difficult due to "domain shift." This paper introduces PCASS, a framework that leverages a simple but effective deep network called PCANet. By explicitly modeling source-specific, target-specific, and shared features—and then aligning them via subspace mapping—the authors achieve superior performance in unsupervised domain adaptation for speech emotion recognition.
The "In-the-Wild" Challenge: Why SER Fails
Most Speech Emotion Recognition (SER) systems operate under a "closed-world" assumption: the training data and test data come from the same distribution. However, in reality, factors like different microphones, ambient noise, and cultural nuances in vocal expression mean that a model trained on one dataset often fails on another.
The core problem is that traditional deep learning is highly sensitive to these perturbations. When labels for the target domain are missing (Unsupervised DA), the model has no "compass" to guide its feature extraction.
Methodology: PCANet Meets Subspace Alignment
The authors' insight is twofold:
- Feature Composition: Common features exist across domains (human emotion is universal), but domain-specific quirks are equally important for classification.
- Filter Alignment: Instead of just transforming data, we should transform the feature extractors (the filters) to point in the "right direction."
1. The PCANet Architecture
Unlike CNNs that learn filters via backpropagation, PCANet uses the principal components of data patches as convolution filters. This makes it efficient and theoretically grounded.
Figure 1: The PCASS framework illustrating the extraction of shared and specific features.
2. The Alignment Mechanism
The PCASS model builds an "interpolating path" between the source and target. It trains three sets of filters: (Source), (Target), and (Shared). To bridge the gap, the source and shared filters are mathematically aligned to the target subspace: This alignment ensures that the resulting features—while still representing the source data—are projected into a space that the target domain "understands."
Experimental Results: Breaking the 60% UAR Barrier
The authors tested their method on the FAU Aibo Emotion Corpus (spontaneous speech from children).
- Baseline (CT): Directly applying a source-trained model yielded only ~51-56% UAR (barely above chance).
- PCASS: Achieved 63.75% with the ABC source and 61.41% with Emo-DB.
Hyperparameter Insight
The study found that larger patch sizes (e.g., 37) are crucial for SER because emotion is not a momentary signal; it spans across longer acoustic temporal windows.
Figure 2: Impact of patch size and filter count on Unweighted Average Recall (UAR) and training time.
Critical Analysis & Takeaways
The brilliance of this work lies in its simplicity. By using PCA-based filters, the authors avoid the volatile training dynamics of typical GAN-based domain adaptation.
Key Takeaways:
- Specifics Matter: Don't just look for "domain-invariant" features. Domain-specific data contains structural information that helps the classifier distinguish between classes.
- Directional Alignment: Aligning the "feature extractors" (filters) is a powerful alternative to aligning the "feature distributions" themselves.
Limitations: While effective, PCANet is a shallow "deep" network. In the era of Transformers, the next step for this research would be applying these subspace alignment principles to the attention heads of a Self-Attention mechanism.
Conclusion
PCASS provides a robust roadmap for unsupervised transfer learning in speech. By treating domain adaptation as a filter-alignment problem rather than just a data-cleaning problem, it sets a high bar for cross-corpus emotion recognition.
