Decoding the Rhythms of Heritage: Machine Learning for Sri Lankan Folk Music Emotion
Machine Learning for Emotion Classification of Sri Lankan Folk Music
This paper introduces a machine learning framework for the automatic emotion classification of Sri Lankan folk music, a culturally unique and previously unexplored dataset. By employing the MATLAB MIRToolbox to extract multi-dimensional acoustic features and testing five standard classifiers, the authors achieved a peak accuracy of 87.95% using the k-Nearest Neighbor (k-NN) algorithm.
TL;DR
While the AI world is obsessed with Western pop, researchers in Sri Lanka are using machine learning to bridge the cultural divide. This paper presents a specialized framework to classify emotions (Happy, Sad, Fear) in Sri Lankan folk melodies, achieving an impressive 87.95% accuracy by optimizing acoustic feature selection.
Background & Motivation: Beyond the Western Canon
Most Music Information Retrieval (MIR) systems are biased toward Western tonality. However, music is not a universal language in the way we often assume—emotional perception is deeply rooted in cultural context. Sri Lankan folk music, tied to traditional livelihoods and innate feelings, offers a rich yet computationally neglected source of emotional expression.
The authors argue that existing classifiers trained on Western datasets lack generalizability. To solve this, they built a native dataset from scratch, involving expert orchestration and verification to ensure "emotionally pure" stimuli.
Methodology: The Anatomy of a Folk Melody
The researchers didn't just throw raw audio at a model; they performed a surgical extraction of musical DNA using the MATLAB MIRToolbox.
The Feature Space
They monitored 22 acoustic features across five perceptual dimensions:
- Dynamics: Energy levels (RMS Energy).
- Rhythm: Tempo, Pulse Clarity, and Matrical Centroids.
- Timbre: The "texture" of the sound (Zero-crossing rate, Brightness, Spectral Flux).
- Pitch: The fundamental frequency and inharmonicity.
- Tonality: Key and tonal centroids.
Classification Workflow
- Pre-processing: Resampling to 44.1kHz and normalization.
- Extraction: Converting 30-second clips into numerical feature vectors.
- Training: Comparing five heavyweights: SVM, Naive Bayes, Decision Tree, Random Forest, and k-NN.
Figure 1: The end-to-end pipeline from raw folk recordings to emotion prediction.
Experiments & Results: The "k-NN" Advantage
Surprisingly, in the context of folk music, the k-Nearest Neighbor (k-NN) algorithm outperformed more complex models like SVM.
- Baseline Performance: k-NN achieved 78.44% accuracy.
- The Breakthrough: By analyzing a Correlation Matrix, the authors discovered that not all 22 features were helpful. Some were noise.
- Optimized Performance: By keeping only the 10 most influential features—specifically those related to brightness, spectral skewness, and pulse clarity—the k-NN accuracy jumped to 87.95%.
Figure 2: Statistical relationships between acoustic features and emotion labels, used to prune the feature set.
Critical Analysis & Conclusion
Takeaway
The success of this study lies in its feature engineering. It proves that for traditional music, "less is more." Focus on the right cultural markers (like rhythmic pulse and spectral brightness) is more effective than sheer model complexity.
Limitations & Future Work
The study currently treats emotion as a single-label problem (one song = one emotion). In reality, folk music often blends complex feelings. Future research should look into:
- Multi-label classification: Handling songs that are both "sad and haunting."
- Deep Learning: Moving from manual feature extraction to raw audio analysis using Deep Neural Networks.
- Expanded Vocab: Including more nuanced emotions like "Calmness" or "Nostalgia."
By digitizing and classifying these melodies, the researchers are not just building a classifier; they are preserving a cultural legacy using the tools of the 21st century.
