"It’s Complicated": Balancing Accessibility with Identity Politics in Image Descriptions
“It’s Complicated”: Negotiating Accessibility and (Mis)Representation in Image Descriptions of Race, Gender, and Disability
The paper investigates how screen reader users, specifically those identifying as BIPOC, non-binary, or transgender, negotiate the description of race, gender, and disability in digital images. It introduces qualitative insights from 25 interviews to inform both human-authored best practices and the ethical development of AI-generated image descriptions.
TL;DR
To a blind person, an image without a description is a locked door; however, an incorrect or reductive AI-generated description can be a form of digital harassment. This paper explores how screen reader users who identify as BIPOC, non-binary, or transgender navigate the high-stakes world of image descriptions—concluding that while more detail is wanted, the power to define one's identity must remain in human hands.
The "Colorblind" Accessibility Trap
For decades, web accessibility guidelines (like WCAG) have been "identity-neutral." The assumption was that describing an image's function was enough. But for a blind person, missing out on the race, gender, or disability status of people in a photo means missing out on the social "cues" that sighted people take for granted.
The authors argue that the current gap in guidelines leads to two failures:
- Erasure: Describers simply ignore appearance to "stay safe," reinforcing a white-cis-normative world.
- Misrepresentation: AI or strangers guess incorrectly, "outing" someone or misgendering them in a way that is persistent and painful.
Methodology: The Nuance of Nonvisual Sensemaking
The research team interviewed 25 screen reader users at the intersection of multiple marginalized identities. They didn't just ask "what do you want to hear?"; they asked "how do you make sense of the world without sight?"
Figure 1: This example illustrates the shift from generic descriptions to identity-affirming ones. Note the use of brackets to signify preferred identity language vs. appearance-only language.
The Critical Distinction: Appearance vs. Identity
The core insight of the paper is the separation of Appearance Phrases and Identity Labels:
- Appearance (The "What"): "Person with darker skin," "short mohawk," "using a white cane."
- Identity (The "Who"): "Black woman," "Non-binary person," "Disabled activist."
Participants generally preferred that strangers/AI stick to appearance (concrete visual facts) while people who actually know the person should use their identity labels.
Results: When Does Appearance Matter?
The study identified six specific contexts where knowing how someone looks is critical for a screen reader user:
- Avatar/Memoji Creation: Trying to make a digital self look like the physical self.
- Encountering New People: Knowing who is talking in a thread or video.
- Identity-Based Discussions: Knowing the "positionality" of participants (e.g., "Is the person speaking about Black issues actually Black?").
- Reading a Room: Determining the safety of a space or finding community members who "look like me."
- Media Representation: Understanding if a movie or brand is truly diverse.
- Product Choices: Knowing if a hair product or clothing item is modeled by someone with a similar body or hair type.
Deep Insight: The AI Dilemma
The study reveals a profound tension regarding Automated Image Description (AID). While AID (like Microsoft's Seeing AI or Facebook's Alt-Text) can scale accessibility, it is often fundamentally reductive.
The Dangers of "Binary" AI
Many participants reported that AI "typically tries to shove you into one characterization or another." For a non-binary person, having an app confidently announce "I see a 40-year-old man" every time they take a selfie is not "access"—it's a microaggression.
Key quantitative takeaway: 10 out of 25 participants were so skeptical of AI bias (citing problems like the "Gender Shades" project) that they advocated against releasing appearance-detection features until better safeguards exist.
Future Outlook: Support, Don't Supplant
The paper concludes with a call for Automation for Support:
- User Control: Let photographees embed their own identity metadata that AI simply "reads" rather than "guesses."
- Context-Awareness: AI should prompt human describers to consider certain visual facts rather than generating the caption autonomously.
- Accountability: Technology providers must be transparent about the "outliers" their models fail to recognize.
Conclusion
Accessibility is not a technical checkbox; it is a social negotiation. As we move toward a world of ubiquitous AI, we must ensure that the "Alt-Text" of the future respects the complexity of the humans it seeks to describe. "Better than nothing" is no longer a high enough bar for AI ethics in accessibility.
