Harmonizing Moods: An End-to-End Model for Emotion-Based Music Selection

Generating Playlists on the Basis of Emotion

2018-11-01
Ganeshsiva Subramaniam, Janhavi Verma, Nikhil Chandrasekhar, Narendra K. C., Koshy George
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an end-to-end framework that generates music playlists based on the emotional state extracted from journal-style text entries. The system integrates a customized text analysis heuristic with an SVM-based music classifier utilizing Spotify API features, achieving optimal performance in mapping text-derived moods to emotionally resonant audio tracks.

TL;DR

Music is the universal language of emotion, yet most streaming services suggest songs based on genre or history rather than your current internal state. This paper bridges that gap by proposing an integrated system that analyzes personal journal entries to extract emotional nuances and recommends a matching playlist. By leveraging the Spotify API and specialized Machine Learning classifiers, the researchers demonstrate a path toward more empathetic digital discovery.

The Missing Link in Affective Computing

Most emotion recognition technology focuses on extrinsic expressions—facial micro-expressions or vocal jitters. However, the written word often captures the intrinsic experience of an emotion. The challenge lies in the pipeline: how do you go from a sentence like "I'm not exactly thrilled" to a specific audio profile?

Prior works often struggled with two main issues:

  1. Textual Nuance: Simple keyword spotting fails to understand that "not happy" is the opposite of "happy."
  2. Class Consistency: Music features (like tempo or loudness) for "Neutral" songs often overlap with "Happy" or "Sad" songs, making algorithmic separation difficult.

Methodology: From Text to Tune

1. The Text Analysis Logic

The authors moved beyond simple polarity (positive/negative) and implemented a customized heuristic. The system generates N-grams (word pairs) to catch "hedge words" (e.g., very, little, not). These are matched against a curated dictionary to calculate an emotional intensity score.

Proposed customized text analysis model

2. Music Feature Extraction

Using the Spotify API, the team extracted six critical dimensions for 579 songs:

  • Valence: Musical positivity.
  • Energy: Intensity and activity.
  • Danceability: Rhythm stability and beat strength.
  • Acousticness, Loudness, and BPM.

3. Finding the Right Classifier

The researchers tested four major ML families: Naïve Bayes, Support Vector Machines (SVM), Decision Trees, and Random Forests.

Experimental Insights

Through rigorous testing on two datasets (Dataset A: curated; Dataset B: volunteer-selected), the results revealed a clear winner.

Classification Results Snippet Excerpt from Table IV: Comparing HAS (Happy-Angry-Sad) performance.

Key Findings:

  • The SVM Edge: The SVM with a linear kernel (SVM-2) was chosen as the optimum model. It maintained high recall across all categories without skewing toward a specific "easy" emotion like Sadness.
  • The "Neutral" Problem: Adding a "Neutral" emotion category significantly dropped accuracy across all classifiers. The audio features of neutral music are often so "spread out" that they blend into other categories, suggesting that neutrality in music is a complex, multivariable state.

Critical Analysis & Real-World Application

The end-to-end test proved the model's intuitive value. When a user input "I am a little melancholy today," the system successfully identified the "Sad" emotion and served tracks like Elvis Presley's "Are You Lonesome Tonight?".

Limitations: The current model relies on a relatively small database of 579 songs. In a real-world scenario with millions of tracks, the "Neutral" overlap would likely worsen without more sophisticated feature engineering (perhaps incorporating lyrics or spectral contrast).

Future Outlook: The paper sets a foundation for "Journal-to-Playlist" features in personal wellness apps. By evolving the psychological model (moving from basic Ekman emotions to more complex dimensions), we could see AI curators that truly understand not just what we like, but how we feel.

Conclusion

This research successfully links two previously parallel aspects of affective computing. By providing a quantitative bridge between linguistic intensity and acoustic signatures, the authors have moved us one step closer to human-centric technology that empathizes with our daily lives.

Takeaway for Devs: When building recommendation engines, don't ignore the "Negation" logic in NLP; it’s the difference between a happy upbeat track and a somber reflection.

Find Similar Papers

Try Our Examples

  • Search for recent papers that improve on Transformer-based emotion detection in long-form journal entries compared to traditional heuristic POS tagging.
  • Which study first defined the relationship between Spotify's 'Valence' and 'Energy' parameters and the Russell circumplex model of affect?
  • Explore how multi-modal emotion recognition (combining text, facial expression, and heart rate) is being used in personalized music therapy or mental health apps.
Contents
Harmonizing Moods: An End-to-End Model for Emotion-Based Music Selection
1. TL;DR
2. The Missing Link in Affective Computing
3. Methodology: From Text to Tune
3.1. 1. The Text Analysis Logic
3.2. 2. Music Feature Extraction
3.3. 3. Finding the Right Classifier
4. Experimental Insights
5. Critical Analysis & Real-World Application
6. Conclusion