Picking Apart the Black Box: Why AI Accessibility is a Sociotechnical Practice, Not a Model Feature
Picking Apart the Black Box: Sociotechnical Contours of Accessibility in AI/ML Software Engineering
This paper presents an ethnographic study conducted at a large global technology firm to examine how AI/ML developers navigate the "black box" nature of their work. It challenges current XAI paradigms by defining accessibility not as a static model feature, but as an emergent sociotechnical practice involving "tinkering" and "situated labor."
TL;DR
Is a model "explainable" simply because it has a heatmap or a SHAP value? This ethnographic study from IBM Research argues no. By following AI/ML developers in the wild, the paper reveals that "access" to these systems is an effortful, ongoing process of human labor. It moves the needle from seeing XAI as a technical module to viewing it as a sociotechnical encounter where developers "tinker" for themselves and "translate" for others.
Contextual Positioning
In the current AI landscape, we are obsessed with "peeking inside the black box" via mathematical interpretability. This paper sits at the intersection of Software Engineering (SE) and Disability Studies, borrowing the concept of "access" to explain how humans gain agency within complex AI ecosystems. It isn’t a paper about a new algorithm; it’s a critique of how we think about the usability of the algorithms we already have.
The Pain Point: The Gap Between XAI and Reality
The author identifies a critical disconnect: XAI researchers build tools to explain models, but in real-world large-scale enterprises (referred to as "TechCorp"), these tools don't automatically make things "accessible."
- Current Failure: We treat "transparency" as a binary—either the box is open or closed.
- The Reality: Different stakeholders need different types of "openness." A developer needs to know about hyperparameters; a doctor needs to know how the model impacts patient outcomes. Treating these as the same problem leads to "glazed eyes" during meetings.
Methodology: The Ethnographic Lens
The researcher engaged in multi-year fieldwork (2017-2019) across four projects.
| Project | Focus | Purpose |
|---|---|---|
| Alpha | Decision-support | Intelligent IT design |
| Beta | Toolkit interface | Evaluative interview study |
| Gamma | Data Science tool | Model improvement practices |
| Delta | NLP Projects | Exploratory interviews on language AI |

Core Insight 1: Tinkering as Sensemaking
For developers, "access" is a process of tinkering.
- They don't just "read" the model; they debug it through surprise. When a model behaves unexpectedly, they seek interpretations.
- Interestingly, developers view current XAI (like heatmaps) as debugging tools for themselves, not as communication tools for end-users. As one informant, Siddhartha, noted: "I can’t see how an end user would find it helpful."
Core Insight 2: Explaining as "Calibration"
The study highlights that explaining AI to others is a creative, iterative "calibration."
- Domain Grounding: Instead of showing equations to medical doctors, successful teams shifted to showing what the model did on specific Electronic Health Records (EHR).
- Reciprocal Literacy: The developer learns the business domain, and the domain expert (SME) learns the model’s boundaries. Access emerges in the middle of this exchange.
Experimental Findings & Visual Evidence
The study highlights the tension between High-Touch interaction (walking users through models) and Scalability.

As projects like Alpha moved toward "General Availability," the dev team couldn't keep doing manual walk-throughs. They tried to "package" accessibility into slide decks and videos, but found that AI's dynamic nature makes these static artifacts obsolete quickly. This suggests a fundamental challenge for SaaS-based AI products.
Critical Analysis & Takeaways
Summary
Accessibility is not a property of the code; it is an interpretive relation. It is something SE teams do, not something a model is.
Limitations
The study is localized to one large global firm ("TechCorp"). While it offers deep qualitative insight, the findings might differ in smaller startups or in consumer-facing AI (B2C) where high-touch "hand-holding" is impossible.
Future Outlook
This work challenges us to stop asking "How do we make this model explainable?" and start asking "How do we design systems that support the collaborative labor of sensemaking?" Future product owners should focus on building tools that facilitate this "calibration" between developers and domain stakeholders rather than just dumping model weights into a UI.
