Beyond Neutrality: Reforming AI Data Curation through Feminist and Race Critical Theory

Ethical Data Curation for AI: An Approach based on Feminist Epistemology and Critical Theories of Race

2021-07-21
Susan Leavy, Eugenia Siapera, Barry O'Sullivan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a critical framework for ethical data curation in AI, grounded in feminist epistemology and critical race theory. It shifts the focus from purely technical bias-mitigation to an interrogation of the social dynamics of power, promoting a "virtue ethics" approach for practitioners to identify and address systemic injustices embedded in machine learning datasets.

TL;DR

The promise of AI as an objective "truth-seeker" is failing because the data it consumes is a reflection of historical power imbalances. This paper argues that "fairness" cannot be solved by math alone. By applying Feminist Epistemology and Critical Race Theory, the authors propose a new framework for data curation that demands an activist stance—moving from a "color-blind" approach to one that actively centers marginalized voices and interrogates the political nature of every data point.

The "Fairness" Trap: Why Mathematical Neutrality is Not Enough

One of the most provocative sections of the paper explores why we can't simply "fix" AI with better algorithms. Using the infamous COMPAS recidivism prediction system as a case study, the authors highlight a fundamental paradox:

  • Calibration: A system can be mathematically "fair" (predicting risk equally across groups).
  • Reality: That same system can stay "calibrated" while simultaneously having a 47% false positive rate for African Americans compared to 24% for Whites.

The paper asserts that because fairness conditions are often mathematically incompatible, developers are always making a political choice. If we ignore the social context (e.g., institutional racism in policing), we aren't being neutral; we are simply reinforcing the status quo.

COMPAS Bias Table Table 1: Disparities in recidivism prediction for African American vs. White defendants illustrate the trade-offs in fairness definitions.

Methodology: The Four Pillars of Ethical Curation

The authors move beyond critique to offer a practical framework for the "Data Curation Professional." The core insight is that knowledge is situated—it always comes from a specific person, place, and power structure.

1. Examining Perspectives

Who labeled the data? Often, the cognitive labor of AI (cleaning and tagging data) is performed by underpaid workers in the Global South, yet the models reflect the "gaze" of white, cis-gendered, male engineers in the Global North. Curation must explicitly document whose world view is being encoded.

2. Analyzing Theory in Data

Is "Race" treated as a biological fact in your dataset, or a social construct? The paper argues that treating identity as a fixed "attribute" risks missing the fluid dynamics of oppression. Curators must analyze the philosophical underpinnings of their features.

3. Reflexivity and Activism

AI doesn't just describe the world; it shapes it through a "runaway feedback loop." The authors call for an anti-racist stance, suggesting that curators shouldn't aim for the "world as it is" (which is biased), but for the "world as it should be" (Data Utopianism).

4. Including Subjugated Knowledges

Traditional datasets often omit "subjugated knowledges"—the lived experiences of those marginalized by history. The solution? Participatory Design.

Ethical Framework Overview The framework emphasizes that data curation is a political act requiring intervention at every stage of the pipeline.

Case Study: The HateTrack Project

To prove this isn't just theory, the paper cites HateTrack, a tool designed to monitor online racist speech. Instead of just scraping keywords, the team:

  • Interviewed members of targeted communities.
  • Had these communities define what "toxicity" looked like to them (e.g., spotting coded language like "religion of peace" used pejoratively).
  • Paid and integrated these experts into the technical team.

The result was a model with a much higher nuance for identifying hate speech that standard, "top-down" models would have classified as neutral.

Critical Insight: The Cost of Participation

The authors are honest about the challenges. Participatory design is "labor-intensive" and can be "traumatizing" for marginalized groups who are asked to repeatedly review toxic content. They warn against "Design Saviors"—engineers who use community knowledge without proper compensation or power-sharing.

Conclusion

This paper is a clarion call for the AI community to stop hiding behind "mathematical objectivity." By embracing feminist and race-critical frameworks, we can transform data curation from a clerical task into a vital act of social justice. The future of "Virtue Ethics" in AI won't be found in better code, but in a more critical, inclusive, and reflexive approach to the data that feeds the machines.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize "Data Feminism" or "Design Justice" frameworks to evaluate bias in Large Language Models (LLMs).
  • Which seminal papers first established the "Impossibility of Fairness" theorem, and how does this paper build upon that mathematical constraint using social theory?
  • Explore how participatory design and "human-in-the-loop" data labeling have been applied specifically to reduce toxicity in multilingual or non-Western AI training sets.
Contents
Beyond Neutrality: Reforming AI Data Curation through Feminist and Race Critical Theory
1. TL;DR
2. The "Fairness" Trap: Why Mathematical Neutrality is Not Enough
3. Methodology: The Four Pillars of Ethical Curation
3.1. 1. Examining Perspectives
3.2. 2. Analyzing Theory in Data
3.3. 3. Reflexivity and Activism
3.4. 4. Including Subjugated Knowledges
4. Case Study: The HateTrack Project
5. Critical Insight: The Cost of Participation
6. Conclusion