Predicting Emotions Perceived from Sounds: A Machine Learning Approach to Sonification

Predicting Emotions Perceived from Sounds

2020-12-10
Faranak Abri, Luis Felipe Gutiérrez, Akbar Siami Namin, David R. W. Sears, Keith S. Jones
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the predictability of emotions perceived from soundscapes using classical machine learning. By utilizing the Emo-Soundscape dataset and MIRToolbox for feature extraction, the authors demonstrate that a Random Forest regression model can effectively predict Arousal and Valence, achieving high accuracy with a reduced set of psychoacoustic features.

TL;DR

Can a machine understand how a sound makes us feel? This research proves it can. By analyzing acoustic features from the Emo-Soundscape dataset, the authors demonstrate that Random Forest models can predict the "Arousal" and "Valence" of sounds with high precision. This paves the way for smarter Cyber-Physical Systems (CPS) that use sound not just to beep, but to communicate intent and emotion.

Background: The Art and Science of Sonification

Sonification is the science of turning data into sound. While we often rely on visualization, sound offers unique advantages in temporal and spatial awareness. However, for a machine to choose the "right" sound for a specific situation (e.g., a critical medical alarm vs. a subtle notification), it must understand the affective impact of that sound.

The study maps emotions onto a 2D space:

  • Arousal: The intensity/excitement level.
  • Valence: The pleasantness (sad to happy).

The Challenge: Why Not Just Use Deep Learning?

Modern AI trends usually point Toward Deep Neural Networks (CNNs/LSTMs). However, this paper takes a refreshing, principled stand on Classical Machine Learning.

  1. Data Scarcity: With only 600 samples, deep models are prone to overfitting.
  2. Independence: Common data augmentation (windowing) risks violating sample independence.
  3. Interpretability: The authors wanted to know which specific acoustic features (like spectral flux or roughness) drive emotional perception.

Methodology: High-Dimensional Acoustic Analysis

The researchers extracted 68 features via the MATLAB MIRToolbox, covering dynamics, rhythm, timbre, and pitch.

Model Architecture and Comparison

They evaluated eight different regressors:

  • Linear Models: Linear Regression, Lasso, ElasticNet, Linear SVR.
  • Non-Linear Models: 2-layer MLP, RBF SVR, Polynomial SVR, and Random Forest.

Russell's Circumplex Model vs Dataset Distribution Figure 1: Scatter plot of Arousal vs. Valence in the Emo-Soundscape dataset, showing a negative correlation (exciting sounds are often perceived as less pleasant).

To refine the models, they used PCA (Principal Component Analysis) and KBest selection to see if a smaller subset of features could maintain accuracy.

Experimental Results: Random Forest Reigns Supreme

The results were conclusive: Random Forest consistently outperformed its peers across all feature sets (All features, PCA-reduced, and Top 25).

Key Performance Metrics (Arousal vs. Valence)

DimensionBest ModelRMSE (Test)R² (Test)
ArousalRandom Forest0.250.84
ValenceRandom Forest0.370.59

Random Forest Hyperparameter Tuning Table 1: Optimal hyperparameters for Arousal and Valence prediction models.

Why is Valence Harder to Predict?

A key insight from the study is that while Arousal is relatively easy to "hear" (loud, sharp, or fast sounds usually mean high arousal), Valence is subjective. Determining if a sound is "pleasant" depends on complex timbral relationships and individual preference, leading to lower R² scores for Valence across all models.

The "Golden" Features

The authors identified 20 core features that are critical for both Arousal and Valence, including:

  • Spectral Flatness: Indicates "noisiness" vs. tone-like quality.
  • Spectral Flux: Measures how quickly the spectrum changes.
  • Spectral Roughness: Related to the sensory dissonance of the sound.

Top Feature Comparison Table 2: The ultimate feature list for emotional prediction.

Conclusion and Future Outlook

This work provides a robust baseline for Automated Sonification. By proving that a manageable set of 25-30 features can predict emotion, the authors enable developers of IoT and Cyber-Physical Systems to select sounds that are not only informative but also emotionally appropriate.

Future work will likely look at Deep Learning as datasets grow, but for now, the "Random Forest on Psychoacoustics" approach offers the best balance of accuracy and interpretability.

Takeaway: When it comes to sound, we don't always need "black box" AI; sometimes, a well-tuned forest of decision trees and a deep understanding of acoustics is all you need to bridge the gap between signal and emotion.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize the Emo-Soundscape dataset with Transformer-based architectures or Attention mechanisms for sound emotion recognition.
  • What are the physiological differences in human response to sound-induced vs. sound-perceived emotions according to recent Affective Computing literature?
  • Which psychoacoustic features are most robust for emotion prediction across different sound taxonomies (e.g., natural sounds vs. mechanical sounds) in IoT applications?
Contents
Predicting Emotions Perceived from Sounds: A Machine Learning Approach to Sonification
1. TL;DR
2. Background: The Art and Science of Sonification
3. The Challenge: Why Not Just Use Deep Learning?
4. Methodology: High-Dimensional Acoustic Analysis
4.1. Model Architecture and Comparison
5. Experimental Results: Random Forest Reigns Supreme
5.1. Key Performance Metrics (Arousal vs. Valence)
5.2. Why is Valence Harder to Predict?
6. The "Golden" Features
7. Conclusion and Future Outlook