Predicting Emotions Perceived from Sounds: A Machine Learning Approach to Sonification
Predicting Emotions Perceived from Sounds
This paper explores the predictability of emotions perceived from soundscapes using classical machine learning. By utilizing the Emo-Soundscape dataset and MIRToolbox for feature extraction, the authors demonstrate that a Random Forest regression model can effectively predict Arousal and Valence, achieving high accuracy with a reduced set of psychoacoustic features.
TL;DR
Can a machine understand how a sound makes us feel? This research proves it can. By analyzing acoustic features from the Emo-Soundscape dataset, the authors demonstrate that Random Forest models can predict the "Arousal" and "Valence" of sounds with high precision. This paves the way for smarter Cyber-Physical Systems (CPS) that use sound not just to beep, but to communicate intent and emotion.
Background: The Art and Science of Sonification
Sonification is the science of turning data into sound. While we often rely on visualization, sound offers unique advantages in temporal and spatial awareness. However, for a machine to choose the "right" sound for a specific situation (e.g., a critical medical alarm vs. a subtle notification), it must understand the affective impact of that sound.
The study maps emotions onto a 2D space:
- Arousal: The intensity/excitement level.
- Valence: The pleasantness (sad to happy).
The Challenge: Why Not Just Use Deep Learning?
Modern AI trends usually point Toward Deep Neural Networks (CNNs/LSTMs). However, this paper takes a refreshing, principled stand on Classical Machine Learning.
- Data Scarcity: With only 600 samples, deep models are prone to overfitting.
- Independence: Common data augmentation (windowing) risks violating sample independence.
- Interpretability: The authors wanted to know which specific acoustic features (like spectral flux or roughness) drive emotional perception.
Methodology: High-Dimensional Acoustic Analysis
The researchers extracted 68 features via the MATLAB MIRToolbox, covering dynamics, rhythm, timbre, and pitch.
Model Architecture and Comparison
They evaluated eight different regressors:
- Linear Models: Linear Regression, Lasso, ElasticNet, Linear SVR.
- Non-Linear Models: 2-layer MLP, RBF SVR, Polynomial SVR, and Random Forest.
Figure 1: Scatter plot of Arousal vs. Valence in the Emo-Soundscape dataset, showing a negative correlation (exciting sounds are often perceived as less pleasant).
To refine the models, they used PCA (Principal Component Analysis) and KBest selection to see if a smaller subset of features could maintain accuracy.
Experimental Results: Random Forest Reigns Supreme
The results were conclusive: Random Forest consistently outperformed its peers across all feature sets (All features, PCA-reduced, and Top 25).
Key Performance Metrics (Arousal vs. Valence)
| Dimension | Best Model | RMSE (Test) | R² (Test) |
|---|---|---|---|
| Arousal | Random Forest | 0.25 | 0.84 |
| Valence | Random Forest | 0.37 | 0.59 |
Table 1: Optimal hyperparameters for Arousal and Valence prediction models.
Why is Valence Harder to Predict?
A key insight from the study is that while Arousal is relatively easy to "hear" (loud, sharp, or fast sounds usually mean high arousal), Valence is subjective. Determining if a sound is "pleasant" depends on complex timbral relationships and individual preference, leading to lower R² scores for Valence across all models.
The "Golden" Features
The authors identified 20 core features that are critical for both Arousal and Valence, including:
- Spectral Flatness: Indicates "noisiness" vs. tone-like quality.
- Spectral Flux: Measures how quickly the spectrum changes.
- Spectral Roughness: Related to the sensory dissonance of the sound.
Table 2: The ultimate feature list for emotional prediction.
Conclusion and Future Outlook
This work provides a robust baseline for Automated Sonification. By proving that a manageable set of 25-30 features can predict emotion, the authors enable developers of IoT and Cyber-Physical Systems to select sounds that are not only informative but also emotionally appropriate.
Future work will likely look at Deep Learning as datasets grow, but for now, the "Random Forest on Psychoacoustics" approach offers the best balance of accuracy and interpretability.
Takeaway: When it comes to sound, we don't always need "black box" AI; sometimes, a well-tuned forest of decision trees and a deep understanding of acoustics is all you need to bridge the gap between signal and emotion.
