Beyond Generic Vectors: Leveraging Controversy for Robust Abusive Language Detection

Twitter-based Polarised Embeddings for Abusive Language Detection

2019-09-01
Leon Graumans, Roy David, Tommaso Caselli
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a method to generate polarised word embeddings by using controversial hashtags on Twitter as proxies for polarized social communities. Using simple Linear SVM classifiers, the authors demonstrate that these embeddings are highly competitive and offer superior cross-domain portability compared to generic embeddings for abusive language detection.

TL;DR

Researchers from the University of Groningen have developed a novel way to detect toxic speech by training "polarised" word embeddings on millions of tweets filtered through controversial hashtags. While standard n-gram models often fail when moving from one platform to another, these polarised embeddings prove remarkably resilient, outperforming generic GloVe vectors in cross-domain scenarios, including technical forums like StackOverflow.

Problem & Motivation: The Portability Gap

Detecting abusive language isn't just about spotting "bad words." It involves understanding the nuances of social interaction, which vary wildly between datasets (e.g., hate speech vs. general offensiveness).

The authors identify a major Achilles' heel in current SOTA-adjacent models: domain sensitivity. A model trained on racism-related tweets often fails to detect sexism or general toxicity because it overfits to specific keywords. Traditional n-gram models are particularly prone to this "lexical bias." The authors' intuition is that by training embeddings on controversial topics—where language is naturally pushed to extremes—they can capture the "flavor" of polarising discourse that leads to abuse, regardless of the specific topic.

Methodology: Engineering Polarization

The core of this work lies in how the data for the embeddings was curated. Unlike Facebook or Reddit, Twitter doesn't have explicit "groups." To overcome this, the authors used 287 controversial keywords as proxies for communities.

The Pipeline:

  1. Keyword Selection: Keywords were pulled from Wikipedia’s list of controversial issues (e.g., abortion, feminism, BLM, Brexit, MAGA).
  2. Data Collection: 6.2 million tweets containing these hashtags were collected (132M tokens).
  3. Embedding Training: Two GloVe models were trained: Polarised 1 (min count 1, window 5) and Polarised 5 (min count 5, window 10).

The "Sanity Check" (Table III) reveals the effectiveness of this approach. While generic embeddings associate "immigrant" with neutral terms like "migrant," the Polarised 5 model pulls in biased terms like "illegal" or "sanctuaries," reflecting the actual linguistic landscape where abuse often occurs.

Table III: Qualitative comparison of nearest neighbors

Experiments: Same Distribution vs. Cross-Domain

The authors tested their models across three major datasets: OffensEval, WH (Waseem & Hovy), and HateEval.

1. In-Domain Results

In a same-distribution scenario, n-gram models still reign supreme. This is expected as they can capture specific derogatory terms used in those specific datasets. However, the polarised embeddings remained competitive, with only a minor performance drop compared to generic GloVe.

2. Cross-Dataset & Cross-Domain (The Real Test)

The true value of polarised embeddings emerged during cross-testing. When a model trained on Twitter was asked to classify "Rude or Offensive" comments on StackOverflow, the polarised embeddings consistently beat the generic ones.

Table VI: Performance on StackOverflow (Out-of-Domain)

The results suggest that Polarised 5 embeddings capture a universal "adversarial" linguistic structure. Whether people are arguing about politics on Twitter or code quality on StackOverflow, the way they express toxicity shares a underlying semantic manifold that polarised embeddings successfully map.

Critical Analysis & Conclusion

Why does it work?

The success of polarised embeddings in the StackOverflow test suggests that controversy is a stylistic marker. By training on controversial hashtags, the model learns the semantics of "us vs. them" narratives, which are common precursors to abusive language.

Limitations

  • Data Volume: The polarised embeddings were trained on significantly less data than generic GloVe Twitter.
  • Indirect Communities: Using focal keywords is a "noisy" proxy for actual communities compared to scraping dedicated hate groups.

Takeaway

This research shifts the focus from what is being said (keywords) to the context of the interaction (polarization). For practitioners building real-world moderation systems, this highlights the necessity of using domain-aware, biased embeddings rather than "clean" generic ones to maintain performance across the shifting sands of social media discourse.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use community-based or user-centric features, such as emojis or interaction networks, to improve the robustness of hate speech detection models.
  • Which study first proposed the use of "controversial topics" for data augmentation in NLP, and how has this evolved with the rise of Transformers and contextualized embeddings like BERT?
  • Explore research that applies polarized word representations to downstream tasks other than abuse detection, such as political sentiment analysis or echo chamber identification.
Contents
Beyond Generic Vectors: Leveraging Controversy for Robust Abusive Language Detection
1. TL;DR
2. Problem & Motivation: The Portability Gap
3. Methodology: Engineering Polarization
3.1. The Pipeline:
4. Experiments: Same Distribution vs. Cross-Domain
4.1. 1. In-Domain Results
4.2. 2. Cross-Dataset & Cross-Domain (The Real Test)
5. Critical Analysis & Conclusion
5.1. Why does it work?
5.2. Limitations
5.3. Takeaway