Beyond Neutrality: Reforming AI Data Curation through Feminist and Race Critical Theory
Ethical Data Curation for AI: An Approach based on Feminist Epistemology and Critical Theories of Race
This paper introduces a critical framework for ethical data curation in AI, grounded in feminist epistemology and critical race theory. It shifts the focus from purely technical bias-mitigation to an interrogation of the social dynamics of power, promoting a "virtue ethics" approach for practitioners to identify and address systemic injustices embedded in machine learning datasets.
TL;DR
The promise of AI as an objective "truth-seeker" is failing because the data it consumes is a reflection of historical power imbalances. This paper argues that "fairness" cannot be solved by math alone. By applying Feminist Epistemology and Critical Race Theory, the authors propose a new framework for data curation that demands an activist stance—moving from a "color-blind" approach to one that actively centers marginalized voices and interrogates the political nature of every data point.
The "Fairness" Trap: Why Mathematical Neutrality is Not Enough
One of the most provocative sections of the paper explores why we can't simply "fix" AI with better algorithms. Using the infamous COMPAS recidivism prediction system as a case study, the authors highlight a fundamental paradox:
- Calibration: A system can be mathematically "fair" (predicting risk equally across groups).
- Reality: That same system can stay "calibrated" while simultaneously having a 47% false positive rate for African Americans compared to 24% for Whites.
The paper asserts that because fairness conditions are often mathematically incompatible, developers are always making a political choice. If we ignore the social context (e.g., institutional racism in policing), we aren't being neutral; we are simply reinforcing the status quo.
Table 1: Disparities in recidivism prediction for African American vs. White defendants illustrate the trade-offs in fairness definitions.
Methodology: The Four Pillars of Ethical Curation
The authors move beyond critique to offer a practical framework for the "Data Curation Professional." The core insight is that knowledge is situated—it always comes from a specific person, place, and power structure.
1. Examining Perspectives
Who labeled the data? Often, the cognitive labor of AI (cleaning and tagging data) is performed by underpaid workers in the Global South, yet the models reflect the "gaze" of white, cis-gendered, male engineers in the Global North. Curation must explicitly document whose world view is being encoded.
2. Analyzing Theory in Data
Is "Race" treated as a biological fact in your dataset, or a social construct? The paper argues that treating identity as a fixed "attribute" risks missing the fluid dynamics of oppression. Curators must analyze the philosophical underpinnings of their features.
3. Reflexivity and Activism
AI doesn't just describe the world; it shapes it through a "runaway feedback loop." The authors call for an anti-racist stance, suggesting that curators shouldn't aim for the "world as it is" (which is biased), but for the "world as it should be" (Data Utopianism).
4. Including Subjugated Knowledges
Traditional datasets often omit "subjugated knowledges"—the lived experiences of those marginalized by history. The solution? Participatory Design.
The framework emphasizes that data curation is a political act requiring intervention at every stage of the pipeline.
Case Study: The HateTrack Project
To prove this isn't just theory, the paper cites HateTrack, a tool designed to monitor online racist speech. Instead of just scraping keywords, the team:
- Interviewed members of targeted communities.
- Had these communities define what "toxicity" looked like to them (e.g., spotting coded language like "religion of peace" used pejoratively).
- Paid and integrated these experts into the technical team.
The result was a model with a much higher nuance for identifying hate speech that standard, "top-down" models would have classified as neutral.
Critical Insight: The Cost of Participation
The authors are honest about the challenges. Participatory design is "labor-intensive" and can be "traumatizing" for marginalized groups who are asked to repeatedly review toxic content. They warn against "Design Saviors"—engineers who use community knowledge without proper compensation or power-sharing.
Conclusion
This paper is a clarion call for the AI community to stop hiding behind "mathematical objectivity." By embracing feminist and race-critical frameworks, we can transform data curation from a clerical task into a vital act of social justice. The future of "Virtue Ethics" in AI won't be found in better code, but in a more critical, inclusive, and reflexive approach to the data that feeds the machines.
