Addressing Bias in Health AI: Why Open Science is the Only Path Forward

Addressing bias in big data and AI for healthcare: a call for open science

2021-05-17
Norori, Natalia, Hu, Qiyang, Aellen, Florence Marcelle, Faraci, Francesca Dalia, Tzovara, Athina
Summary
Problem
Method
Results
Takeaways
Abstract

This perspective paper identifies the multidimensional sources of bias in medical AI—categorized into data-driven, algorithmic, and human factors—and proposes a strategic framework rooted in Open Science to foster fairness. The authors advocate for participant-centered development, inclusive data standards, and code sharing as essential components to prevent AI from reinforcing historical healthcare inequalities.

Executive Summary

TL;DR: AI has the potential to revolutionize healthcare, but it is currently a "double-edged sword." Due to deep-seated biases in data collection and algorithm design, AI systems frequently fail the very populations that need them most. This paper argues that Open Science—characterized by data transparency, participatory design, and code sharing—is the most robust framework for identifying and dismantling these biases.

Context: This is a foundational Perspective piece that moves beyond merely identifying bias; it provides a roadmap for researchers to transition from "Black Box" healthcare to an inclusive, accountable, and "White Box" (explainable) clinical future.

The Anatomy of a Medical Crisis: Where Bias Hides

The authors dissect bias into three critical vectors that frequently overlap:

  1. Data-Driven Bias: Most datasets are "WEIRD" (Western, Educated, Industrialized, Rich, and Democratic). If a model only sees white skin during training, its diagnostic accuracy for melanoma drops by half when facing Black patients.
  2. Algorithmic Bias: Standard optimization metrics (like Accuracy) are deceptive. In a dataset that is 80% healthy, a model can ignore the minority class entirely and still "score" 80%.
  3. Human Bias: Engineers often solve problems they personally relate to, inadvertently ignoring the needs of the elderly, immigrants, or the LGBTQ+ community.

Sources of Bias in AI Figure 1: Illustration of diverse sources of bias in machine learning, ranging from historical prejudices to technical data gaps.

Methodology: The Open Science Toolkit

The core of the paper is the "Call for Open Science." The authors argue that the traditional closed-off nature of medical research is what allows bias to persist. They propose a transition to:

1. Participatory Science

Instead of developing for patients, developers must work with them. The OpenAPS (Open Artificial Pancreas) project is cited as a gold standard—a patient-led initiative where the target community actually wrote the code and managed the data.

2. Algorithmic Transparency & Explainability

The authors emphasize the switch to Explainable AI (XAI). Using techniques like Grad-CAM, clinicians can see which pixels a network is looking at. This prevents the model from making a "right" diagnosis for the "wrong" reason (e.g., identifying a surgical tool instead of a lesion).

3. Synthetic Data & GANs

When real-world clinical data is scarce for minority groups, Generative Adversarial Networks (GANs) can synthesize realistic "underrepresented data" to balance training sets without compromising privacy.

Open Science Tools for Fairness Figure 2: The Open Science framework for addressing bias through data sharing, standard interoperability, and code transparency.

Case Study: The Pulse Oximeter Failure

The paper opens with a haunting example: the pulse oximeter. Because it was developed primarily on light-skinned individuals, it systematically overestimates oxygen levels in non-white patients. Black patients are three times more likely to have undetected low blood oxygen (occult hypoxemia). This isn't just a technical glitch; it is a life-threatening failure of design that Open Science could have exposed during the prototyping phase.

Critical Insight & Conclusion

The authors conclude that "invisible" data leads to "invisible" patients. If healthcare AI is to be integrated into clinical routines, it must undergo the same rigorous, transparent testing as clinical trials.

Takeaways for the Industry:

  • Move beyond Accuracy: Use the F1-Score and class-weighted loss functions to handle imbalanced datasets.
  • Incentivize Transparency: Preregister AI studies to prevent "p-hacking" or results-doctoring.
  • Privacy is not an Excuse: Use Federated Learning to train on diverse local hospital data without moving sensitive patient records.

Limitations: While Open Science is a powerful tool, it requires a global shift in institutional policy and funding—obstacles that are often more political than technical.


Summary of "Addressing bias in big data and AI for health care: A call for open science" (Norori et al., 2021).

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Generative Adversarial Networks (GANs) specifically to augment medical imaging datasets for underrepresented ethnic groups.
  • Which original paper established the "WEIRD" population bias framework, and how has this concept evolved in the context of Big Data and AI ethics?
  • Find research implementations of Federated Learning in multi-hospital settings that specifically measure the reduction of algorithmic bias in diagnostic accuracy.
Contents
Addressing Bias in Health AI: Why Open Science is the Only Path Forward
1. Executive Summary
2. The Anatomy of a Medical Crisis: Where Bias Hides
3. Methodology: The Open Science Toolkit
3.1. 1. Participatory Science
3.2. 2. Algorithmic Transparency & Explainability
3.3. 3. Synthetic Data & GANs
4. Case Study: The Pulse Oximeter Failure
5. Critical Insight & Conclusion