Beyond Keywords: Bridging Psychology and AI to Solve Online Conflictual Language

Automatic Identification of Harmful, Aggressive, Abusive, and Offensive Language on the Web: A Survey of Technical Biases Informed by Psychology Literature

2021-09-30
Agathe Balayn, Jie Yang, Zoltán Szlávik, Alessandro Bozzon
Summary
Problem
Method
Results
Takeaways

This paper presents a comprehensive survey of Online Conflictual Language (OCL) detection, proposing a reconciled taxonomy informed by psychology literature. It systematically maps technical gaps in current NLP systems and provides a framework for addressing biases in dataset engineering and machine learning models.

TL;DR

The automatic detection of hate speech and cyberbullying is often treated as a solved problem in "leaderboard" culture, but it remains broken in the real world. This survey by Balayn et al. highlights a massive terminological and conceptual mismatch between how computer scientists build models and how humans experience conflict. By introducing a psychology-informed taxonomy, the authors expose why current SOTA models fail to generalize and offer a roadmap for building more robust, fair, and context-aware moderation systems.

The "Decontextualization" Crisis

Most current NLP models for abusive language operate in a vacuum. A model sees a string of text, checks for "toxic" tokens, and spits out a probability.

Why is this a problem? Psychology tells us that the perception of harm is inherently subjective. A phrase used jokingly between friends (context) is vastly different from the same phrase used as a slur by a stranger. Current technical pipelines ignore this by:

  • Uniforming subjective labels: Using majority voting to crush dissenting annotator opinions, which often silences the perspectives of marginalized groups.
  • Ignoring the Observer: Failing to account for the fact that a woman might perceive a tweet as more aggressive than a man would, based on lived experience.

Methodology: A Multi-Dimensional Taxonomy

The core contribution of this work is a reconciled taxonomy that moves beyond "Hate Speech" as a catch-all term. The authors define Online Conflictual Language (OCL) through seven distinct anchors.

Table 8: Detailed Overview of OCL Individual Characteristics

This framework allows us to distinguish between:

  1. Aggression: Driven by parental/behavioral intent to harm.
  2. Offensive Language: Centered on the target's characteristics.
  3. Abusive Language: Defined by the style (e.g., profanity) rather than the target.

Technical Biases in the Pipeline

The paper systematically deconstructs the machine learning pipeline to find where "biases" creep in:

1. Data Collection (The Retrieval Bias)

Most datasets are built using keyword-based sampling. If you only collect data containing "bad words," your model becomes a glorified profanity filter. It will fail to detect "coded" hate speech or microaggressions that use polite but exclusionary language.

2. Annotation (The "Majority Rule" Bias)

Current practices favor Fleiss’ Kappa or high agreement metrics. However, for OCL, disagreement is often data, not noise. By forcing a single ground truth, we introduce Aggregation Bias, where the "average" (often majority-group) opinion becomes the model's objective reality.

3. Feature Engineering (The Context Mismatch)

While deep learning (CNNs/RNNs) has improved accuracy, most models still only look at the textual content. As shown in the study, only a fraction of papers utilize metadata like user history, network centrality, or conversation threading.

Type of Information Used by Classification Methods

Critical Insight: The "Generalization" Wall

One of the most sobering results discussed is the performance drop when models move from "Laboratory" datasets to "Deployment" data. A model achieving a 70 F1-score on its native test set can drop to a mere 21.1 F1 when tested on a different platform. This is a direct result of the technical biases listed above—the model learns the specific quirks of the dataset rather than the general nature of conflict.

Future Outlook: Building Truly Mature Systems

To move forward, the authors advocate for:

  • Human-in-the-Loop 2.0: Using "CrowdTruth" approaches that embrace disagreement.
  • Counter-speech Generation: Moving from simple "delete" moderation to more sophisticated linguistic interventions.
  • Intersectional Evaluation: Ensuring models don't just work "on average" but are fair across different demographic slices (Gender, Ethnicity, Age).

Conclusion

Balayn et al. remind us that OCL detection is not just a "math problem" to be optimized. It is a social science challenge that requires technical humility. We cannot build safer online spaces until our models understand the Why and the Who, not just the What.

Find Similar Papers

Try Our Examples

  • Search for recent papers attempting to solve the problem of subjectivity and inter-annotator disagreement in toxic comment classification using non-majority voting methods.
  • Which paper first proposed the use of "Buss-Perry Aggression Questionnaire" (BPAQ) in computational linguistics, and how have subsequent works adapted psychological scales for social media?
  • Explore research that applies the authors' proposed multidimensional OCL taxonomy to multimodal detection tasks involving images and memes.
Contents
Beyond Keywords: Bridging Psychology and AI to Solve Online Conflictual Language
1. TL;DR
2. The "Decontextualization" Crisis
3. Methodology: A Multi-Dimensional Taxonomy
4. Technical Biases in the Pipeline
4.1. 1. Data Collection (The Retrieval Bias)
4.2. 2. Annotation (The "Majority Rule" Bias)
4.3. 3. Feature Engineering (The Context Mismatch)
5. Critical Insight: The "Generalization" Wall
6. Future Outlook: Building Truly Mature Systems
6.1. Conclusion