WASABI: Bridging the Gap Between Cultural Metadata and Multimodal Music Analysis

The WASABI Dataset: Cultural, Lyrics and Audio Analysis Metadata About 2 Million Popular Commercially Released Songs

2021-01-01
Michel Buffa, Elena Cabrio, Michael Fell, Fabien Gandon, Alain Giboin, Romain Hennequin, Franck Michel, Johan Pauwels, Guillaume Pellerin, Maroua Tikat, Marco Winckler
Summary
Problem
Method
Results
Takeaways
Abstract

The WASABI dataset is a massive-scale Knowledge Graph encompassing over 2 million popular songs, 200k albums, and 77k artists. It uniquely integrates cultural metadata with high-level features derived from multimodal analysis of both lyrics (NLP) and audio signals (MIR).

TL;DR

The WASABI project has unveiled a monumental dataset of 2 million commercial songs, transforming the landscape of Music Information Retrieval (MIR). Unlike previous datasets that focus solely on audio or metadata, WASABI builds a Semantic Knowledge Graph that connects artist biographies and discographies with deep NLP-based lyrics analysis (emotions, topics) and audio features (chords, tempo). It provides a unified, open-source framework for musicologists, journalists, and developers.

The "Missing Link" in Music Data

For years, the music research community faced a trade-off. You could have the depth of cultural metadata (who produced the song? what genre?) from sources like MusicBrainz, or the breadth of raw audio signals from the Million Song Dataset (MSD). However, finding a dataset that tells you why a song feels sad through its lyrics while proving it with its chord progression was nearly impossible at scale.

The authors argue that the "heart" of a song lies in the intersection of its cultural context, lyrical storytelling, and sonic structure. The WASABI project was born to fill this gap, doubling the size of MSD and adding a sophisticated semantic layer.

Methodology: The Multimodal Pipeline

The construction of WASABI is a masterclass in data engineering and machine learning integration.

1. The Knowledge Graph Architecture

The core of WASABI is its RDF (Resource Description Framework) Knowledge Graph. By extending the standard Music Ontology, the team created a format where a song is not just a file name, but a node in a web of relationships:

  • Cultural Data: Merged from LyricsWikia, DBpedia, MusicBrainz, and Discogs.
  • Lyrics Analysis: BERT-based emotion detection (Valence-Arousal), LDA for topic modeling, and CNNs for structural segmentation (identifying Verse vs. Chorus).
  • Audio Analysis: Automatic chord recognition with confidence scores and integration with the TimeSide API for real-time signal processing.

WASABI Pipeline Figure 1: The WASABI pipeline showing the flow from raw data sources to the synchronized Knowledge Graph and end-user applications.

2. Semantic Enrichment of Lyrics

WASABI doesn't just store lyrics; it understands them. Using a Convolutional Neural Network (CNN) trained on self-similarity matrices, the system can detect the repetitive structure of a song. Combined with BERT for emotion classification, it categorizes how "joyful" or "tragic" a track is based on its text—a feature highly valuable for sentiment-based recommendation.

Experimental Results & Data Quality

Building a dataset of 2 million entries from crowd-sourced wikis is inherently noisy (e.g., spelling variations like "Omega Man" vs. "Ω Man"). To ensure SOTA quality, the authors conducted:

  • WASABI Marathons: Human-in-the-loop hackathons to identify and fix metadata conflicts.
  • Cross-Source Matching: Successfully linking 78% of artists and 72% of songs to established external authorities (MusicBrainz/Deezer).

Dataset Statistics Table 1: Volume of data processed, resulting in over 55 million RDF triples for the community.

Impact: More Than Just a Database

The utility of WASABI extends beyond academic research. The authors highlight several real-world scenarios:

  • Journalism: Radio presenters searching for "protest songs" or "anti-government themes" during social movements.
  • Music Education: Students analyzing the chord complexity of the Blues in the key of E across different decades.
  • Recommendation: Moving beyond "Users who liked X also liked Y" to "Find songs with similar emotional arcs and lyrical topics."

Critical Insight & Limitations

While WASABI is a breakthrough, it faces a significant hurdle: Copyright. Because lyrics and audio are proprietary, the authors can only provide the metadata and the analysis results (like chord symbols without timing). Researchers wishing to reproduce the work still need to interface with commercial APIs (like MusixMatch or Deezer) to fetch the raw content.

However, as an open-source "anchor" in the Linked Data cloud, WASABI provides the necessary identifiers (IDs) to link these disparate sources once the user has the appropriate permissions.

Conclusion

The WASABI dataset is more than a collection of songs; it is a specialized brain for popular music. By formalizing the relationship between the "vibe" (audio), the "story" (lyrics), and the "heritage" (metadata), it paves the way for a more intelligent, semantic understanding of our musical culture.


For more details, check out the WASABI Explorer and their GitHub Repository.

Find Similar Papers

Try Our Examples

  • Look for recent papers that utilize the WASABI dataset for multi-modal music recommendation or emotion-based playlist generation.
  • Which studies first established the Music Ontology (MO) framework, and how does the WASABI ontology extension specifically handle lyrics-related semantic properties?
  • Investigate how deep learning models like BERT and CNNs are currently state-of-the-art for lyric segmentation compared to traditional signal processing methods.
Contents
WASABI: Bridging the Gap Between Cultural Metadata and Multimodal Music Analysis
1. TL;DR
2. The "Missing Link" in Music Data
3. Methodology: The Multimodal Pipeline
3.1. 1. The Knowledge Graph Architecture
3.2. 2. Semantic Enrichment of Lyrics
4. Experimental Results & Data Quality
5. Impact: More Than Just a Database
6. Critical Insight & Limitations
7. Conclusion