HateClassify: Rethinking Hate Speech Detection as a Multi-Label Challenge

HateClassify: A Service Framework for Hate Speech Identification on Social Media

2020-11-10
Muhammad Usman Shahid Khan, Assad Abbas, Attiqa Rehman, Raheel Nawaz
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "HateClassify," a CNN-based service framework for social media hate speech identification. It redefines hate speech detection as a multi-label classification problem rather than a traditional multi-class one, utilizing a Sequential Convolutional Neural Network (SCNN) to outperform baseline models.

TL;DR

The explosion of toxic content on social media (a 900% increase between 2014 and 2016) has made automated moderation critical. HateClassify is a new framework that shifts the paradigm from "Is this A or B?" to "To what degree is this A and B?". By treating hate speech as a multi-label problem and employing a Sequential CNN (SCNN), researchers achieved a 20% boost in detection performance compared to traditional methods.

The "Hate vs. Offensive" Paradox

Social media companies like Twitter and Facebook struggle with a fine line: what separates a vulgar insult from actual hate speech?

The authors identify a critical technical flaw in existing research: vocabulary overlap. In some datasets, 65.2% of words used in "hate" tweets are also present in "offensive" tweets. When we force a machine (or a human) to pick just one label, the nuance is lost. This ambiguity causes traditional multi-class models to suffer from high error rates in minority classes.

Methodology: The HateClassify Framework

The proposed framework consists of two main pillars: a democratic crowd-sourced labeling policy and a robust SCNN model.

1. The Crowd-Sourced Service Model

Instead of a top-down approach where a corporation defines hate speech globally, HateClassify proposes a service where local users vote on content. This allows for geographical sensitivity—recognizing that what is offensive in one cultural context might be legal or culturally different in another.

2. The SCNN Architecture

The core of the detection engine is a Sequential Convolutional Neural Network (SCNN).

  • Input: Word embeddings (256-dimensional).
  • Layers: Three 1D Convolutional layers with kernel sizes of 3, 4, and 5 to capture varying n-gram context.
  • Mechanism: Unlike standard classifiers, it uses a Sigmoid activation in the final dense layer with an alpha-evaluation metric to allow for multiple labels per tweet.

HateClassify Framework Figure 1: The proposed service framework showing the offline training and online detection modules.

Experimental Insights

The researchers tested SCNN against strong baselines, including SVMs and Attentive CNNs (ATTCNN), across three major datasets.

Visualizing the Problem: ScatterText

By using ScatterText, the authors proved that "Hate" and "Offensive" categories are spatially inseparable in many cases, justifying the need for multi-label logic.

ScatterText Dataset 2 Figure 2: ScatterText visualization showing the high frequency of shared words between hate and offensive labels.

Results: The Power of Multi-Label

The move to multi-labeling wasn't just a conceptual shift—it delivered massive performance gains:

  • 20% Increase in hate speech detection performance.
  • Higher Precision: SCNN was more "stringent" than Logistic Regression, meaning it had fewer false positives when identifying the most toxic content.
  • Robustness: SCNN maintained stable performance even when datasets were highly unbalanced (e.g., in the Sexism vs. Racism dataset).

Critical Analysis & Future Outlook

While SCNN proved powerful, the study noted that "Attention Mechanisms" (ATTCNN) actually decreased recall in this specific task by being too selective. This suggests that for short-form text like tweets, simple sequential convolution might be more effective than complex attention modules.

Takeaway: The future of content moderation lies in probabilistic labeling. By acknowledging that a post can be "offensive" AND "hateful" simultaneously, we build systems that better reflect the messiness of human language.

Limitations: The framework currently relies heavily on manual crowd-sourced labels for retraining. Integrating this with unsupervised pre-training or LLM-assisted labeling could significantly scale the system's efficiency.

Find Similar Papers

Try Our Examples

  • Which recent papers explore the use of Multi-Label Learning (MLL) specifically to resolve label ambiguity in toxic comment classification on social media?
  • What is the theoretical origin of the alpha-evaluation metric for multi-label classification, and how do modern variants handle severe class imbalance?
  • How can the crowd-sourced "HateClassify" framework be integrated with Large Language Models (LLMs) to provide real-time, explainable content moderation?
Contents
HateClassify: Rethinking Hate Speech Detection as a Multi-Label Challenge
1. TL;DR
2. The "Hate vs. Offensive" Paradox
3. Methodology: The HateClassify Framework
3.1. 1. The Crowd-Sourced Service Model
3.2. 2. The SCNN Architecture
4. Experimental Insights
4.1. Visualizing the Problem: ScatterText
4.2. Results: The Power of Multi-Label
5. Critical Analysis & Future Outlook