Wiki-MID: Bridging the Cross-Domain Gap with Twitter and Wikipedia

Wiki-MID: A Very Large Multi-domain Interests Dataset of Twitter Users with Mappings to Wikipedia

2018-01-01
Giorgia Di Tommaso, Stefano Faralli, Giovanni Stilo, Paola Velardi
Summary
Problem
Method
Results
Takeaways
Abstract

Wiki-MID is a very large, LOD-compliant multi-domain interests dataset of Twitter users (English and Italian) mapped to Wikipedia. It includes approximately 500,000 users with an average of 90 preferences per user across domains such as music, books, movies, and politics, achieving state-of-the-art scale for cross-domain recommendation research.

TL;DR

Researchers from Sapienza University of Rome have released Wiki-MID, a massive Linked Open Data (LOD) compliant dataset that mappings nearly half a million Twitter users' multi-domain interests to Wikipedia entries. By combining explicit sharing behavior (e.g., Spotify #NowPlaying) with implicit social graph analysis, the dataset provides a high-density, high-precision resource for training the next generation of semantic recommender systems.

The "Data Silo" Problem in Recommendations

Modern Recommender Systems (RS) are only as good as the data they eat. However, the academic community faces a major hurdle: domain isolation. Most public datasets like MovieLens or the Million Song Dataset are "single-note" performers. Real-world users, however, are multi-faceted; they read books on Goodreads, listen to music on Spotify, and discuss politics on Twitter.

Building a multi-domain view is traditionally the privilege of "Big Tech" giants. For independent researchers, the challenge is twofold:

  1. Data Sparsity: Users rarely link their accounts across multiple niche platforms.
  2. Semantic Ambiguity: "High" could be a song title, a movie, or a book—without context, the data is noise.

Methodology: The Hybrid Extraction Strategy

The authors utilize a sophisticated three-step workflow to turn noisy Twitter streams into a structured interest graph.

1. Explicit Interests (The "Service" Path)

The system monitors hashtags like #NowPlaying, #IMDb, and #Goodreads. Instead of relying on the text of the tweet (which is often ungrammatical or ambiguous), the methodology scrapes the URL within the tweet. This provides clean metadata (Title, Author, Year) directly from the source platform.

2. Implicit Interests (The "Topical Friend" Path)

This is the paper's most significant contribution to scale. The authors define Topical Friends as followees who represent an interest rather than a social peer.

  • The Classifier: They trained an SVM model using features like in-degree/out-degree ratios and profile keywords (e.g., "artist", "politician") to distinguish between a user's real-life friend and a celebrity/brand account.

3. Wikipedia Grounding (The Ensemble Mapper)

To transform a Twitter handle (like @nytimes) into a global concept, the authors use an ensemble of three mapping techniques:

  • Context-based: Using BabelNet for BoW (Bag-of-Words) similarity.
  • Link-based (M2/M3): Matching URLs from Twitter profiles to DBpedia homepage properties.

Overall Methodology and Data Model Figure 1: The data model using SIOC and SKOS ontologies to link user accounts to Wikipedia entities.

Experimental Results & Dataset Quality

The scale of Wiki-MID is objectively impressive. For the English stream alone:

  • Users: 444,744
  • Unique Interests: ~341,000 mapped to Wikipedia.
  • Precision: 90% - 96% verified by human annotators.

The "Topical Friends" method significantly boosts the density of the dataset. As shown in the figure below, the distribution of interests per user peaks much higher than traditional datasets, with 100,000 users possessing over 100 distinct interests.

Interest Distribution Figure 2: Distribution of interests induced from topical friends (English left, Italian right).

Compared to previous benchmarks (like Dooms et al.), which struggled to find users with multiple domains, Wiki-MID provides a rich "Venn Diagram" of overlapping interests in music, movies, and books.

Critical Insights: Why it Matters

The true value of Wiki-MID isn't just its size—it's the Wikipedia Mapping. By tethering a user's interest to a Wikipedia URI, researchers can leverage the entire LOD (Linked Open Data) Cloud.

  • Generalization: If a user likes "The Magnetic Fields," the system can look at the Wikipedia Category Graph to see they are associated with "Indie Pop."
  • Cross-Lingual Capability: Because Wikipedia IDs are language-independent, a model trained on Italian users could potentially generalize to English preferences.

Conclusion & Limitations

Wiki-MID is a masterclass in "data alchemy," turning the chaotic lead of Twitter streams into the gold of structured semantic data. However, it relies on unary ratings (i.e., we know they like something, but not how much compared to something else).

For researchers looking to solve the "cold start" problem in cross-domain recommendation, Wiki-MID is likely the most comprehensive open-source playground currently available.


Editor's Note: The dataset and software are released under Creative Commons, providing a vital resource for the RecSys community to benchmark semantic algorithms.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Wikipedia Category Graphs or DBpedia for cross-domain recommendation systems since 2018.
  • Who first proposed the concept of "topical friends" in social network analysis, and how has the definition evolved in the context of user profiling?
  • Find research that applies the Semantically-Interlinked Online Communities (SIOC) ontology to modern Large Language Model (LLM) user modeling.
Contents
Wiki-MID: Bridging the Cross-Domain Gap with Twitter and Wikipedia
1. TL;DR
2. The "Data Silo" Problem in Recommendations
3. Methodology: The Hybrid Extraction Strategy
3.1. 1. Explicit Interests (The "Service" Path)
3.2. 2. Implicit Interests (The "Topical Friend" Path)
3.3. 3. Wikipedia Grounding (The Ensemble Mapper)
4. Experimental Results & Dataset Quality
5. Critical Insights: Why it Matters
6. Conclusion & Limitations