EDA-Optimized Emotion Recognition: Lifting the Language Barrier in Spanish and Basque
Feature Subset Selection Based on Evolutionary Algorithms for Automatic Emotion Recognition in Spoken Spanish and Standard Basque Language
This paper presents a framework for automatic emotion recognition in Spoken Spanish and Standard Basque using the RekEmozio bilingual database. The authors employ Feature Subset Selection (FSS) based on Estimation of Distribution Algorithms (EDA) to optimize the performance of various Machine Learning classifiers across seven emotional states.
TL;DR
Researchers have developed a more accurate way to detect human emotions in spoken Spanish and Standard Basque. By applying Estimation of Distribution Algorithms (EDA) for feature selection, the study demonstrated a massive accuracy jump (over 15%) compared to standard methods, proving that selecting the right "voice fingerprints" is more important than simply having more data.
Problem & Motivation
The challenge of Affective Computing lies in the nuance of the human voice. While significant research exists for English-language emotion recognition, bilingual and regional contexts—such as the interplay between Spanish and Standard Basque—are often overlooked.
Current systems frequently face the "curse of dimensionality." When extracting 32 global paralinguistic parameters (like pitch, jitter, and energy spectral distribution), many features are redundant or introduce noise. Standard Machine Learning (ML) models like Decision Trees or Naive Bayes often "get confused" by this irrelevant data, leading to mediocre recognition rates.
Methodology: The Power of Evolutionary Selection
The core innovation of this work is the application of Feature Subset Selection (FSS) using a wrapper-based Estimation of Distribution Algorithm (EDA).
1. Feature Extraction (The Raw Material)
The team extracted 32 acoustic features from the RekEmozio database, categorized into:
- Prosodic Features: F0 (Pitch) statistics and RMS Energy.
- Voice Quality: Jitter (frequency perturbation) and Shimmer (amplitude perturbation).
- Spectral Features: Energy distribution across low, mid, and high bands.
- Temporal Features: Speaking rate and silence durations.
2. The EDA Wrapper
Instead of using all 32 features, the EDA treats feature selection as an optimization problem. It builds a probabilistic model of the most "successful" feature combinations.
- The Loop: Pick a subset -> Train the classifier (IB, C4.5, NB) -> Evaluate via 10-fold cross-validation -> Update the probability of choosing those features.
- Why it works: Unlike "greedy" algorithms that might get stuck in local optima, EDA explores the search space more effectively by learning the dependencies between features.
(The Naive Bayes rule serves as one of the baseline comparison paradigms used within the feature evaluation framework.)
Experimental Results: Quality Over Quantity
The results were striking. In the baseline experiments using all variables, the performance was lackluster (averaging around 41-48%). However, once the evolutionary FSS was applied, the Instance-Based (IB) learning paradigm saw a dramatic surge.
Key Performance Gains:
- Standard Basque: Accuracy rose from ~41% to 64.88%.
- Spanish: Accuracy rose from ~38-45% to ~58-68% (depending on gender/classifier).
- The Winner: The IB (Instance-Based) classifier consistently outperformed Decision Trees (C4.5/ID3) once the feature set was pruned.
(Table 3: Accuracy after FSS optimization specifically for Basque, showing the dominance of the IB classifier across male and female speakers.)
Critical Analysis & Conclusion
Takeaway
The success of this study underscores that in paralinguistic analysis, Inductive Bias (the assumptions a model makes) is heavily influenced by the feature set. By removing the "noise" of irrelevant features through evolutionary algorithms, even simple classifiers like k-NN (IB) can outperform complex decision trees.
Limitations & Future Work
While the 15% improvement is a breakthrough for these specific languages, total accuracy remains in the 60% range. The authors suggest that the next frontier is Multimodal Fusion—combining these optimized audio features with facial gesture recognition and physiological data to reach the high-fidelity recognition required for real-world AI assistants.
This research provides a vital blueprint for building affective systems in "under-represented" linguistic regions, ensuring that AI can understand not just what is said, but how it is felt, across different cultures.
