Crowdsourcing the Kidney: Scalable Medical Annotation Without the Expert Price Tag

Segmenting The Kidney On CT Scans Via Crowdsourcing

2019-04-01
Paras Mehta, Veit Sandfort, Daan Gheysens, Gert-Jan Braeckevelt, Jonathan Berte, Ronald M. Summers
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the feasibility of using crowdsourcing from untrained workers to annotate kidney segmentations on CT scans for training Deep Learning models. By employing a majority-voting consensus from 72 users on the Robovision AI platform, the authors demonstrated that non-expert annotations can train a 3D U-Net to achieve SOTA-level performance (Dice Score = 0.904), comparable to models trained on expert data.

TL;DR

Training robust Convolutional Neural Networks (CNNs) for medical imaging is notoriously data-hungry and expensive. This study proves that untrained crowd workers, paired with a simple majority-voting system, can generate kidney segmentations on CT scans that are as effective for training AI models as those produced by medical experts. The resulting model achieved a 0.904 Dice score, rivaling expert-trained performance while drastically reducing time and cost.

Problem & Motivation: The Expert Bottleneck

The "AI revolution" in radiology faces a major hurdle: The Annotation Bottleneck. Unlike general computer vision (where anyone can identify a "cat" or "dog"), medical image segmentation—delineating an organ's borders in 3D—requires specialized knowledge.

Conventionally, this task falls on radiologists, whose time is both expensive and scarce. Previous SOTA models for tasks like diabetic retinopathy have required over a million images. For 3D CT scans, where each organ must be traced through dozens of axial slices, the labor cost is astronomical. The authors asked a bold question: Can we use the "wisdom of the crowd" to replace the precision of the expert?

Methodology: From Polygons to Consensus

The researchers used 42 CT scans from the NIH "Pancreas-CT" dataset. However, instead of asking a radiologist to spend 37 minutes per scan, they offloaded the task to 72 untrained workers via the Robovision AI (RVAI) platform.

The Workflow:

  1. Training: Workers completed a short module to recognize the kidney's position relative to the spine.
  2. Annotation: 5 to 10 qualified users drew polygons around the kidneys in 2D slices.
  3. Majority Voting: To eliminate individual noise, the team used a pixel-level vote. A pixel was labeled "kidney" only if >70% of workers agreed.
  4. Interpolation: Since only every 5th slice was labeled, 3D volumes were reconstructed using interpolation.

The RVAI Platform for Crowdsourced Segmentation Figure 1: The interface where untrained users perform polygon-based organ segmentation.

The Post-Processing Pitfall:

Interestingly, the authors tried using the GrabCut algorithm to "clean up" the human annotations. Surprisingly, this decreased accuracy (Dice score dropped from 0.938 to 0.895). The human workers were better at distinguishing the kidney from adjacent renal vessels than the automated refinement algorithm was.

Experiments & Results: Crowd vs. Expert

The authors trained a 3D U-Net on three different data sources. The performance metric used was the Dice Similarity Coefficient (DSC), where 1.0 represents a perfect overlap.

Training SetTest Dice ScoreStatistical Significance
Expert Labeled0.885 ± 0.112Baseline
Crowdsourced0.904 ± 0.026P = 0.50 (No significant difference)
Combined (Expert + Crowd)0.932 ± 0.040Best overall performance

Performance Validation Graphs Figure 2: Statistical comparison showing that crowd-labeled 3D scans (C) achieved high similarity to reference standards.

The results were clear: The CNN trained on the "crowd" was not just "good enough"—it was statistically indistinguishable from the expert-trained model. Furthermore, the combined training set (Set A) showed that adding crowd data to expert data improves model stability and reduces standard deviation.

Critical Analysis & Conclusion

Takeaway

The "Expert-in-the-loop" model is moving toward an "Expert-as-Supervisor" model. This study demonstrates that for distinct anatomical structures like the kidney, crowdsourcing provides a massive throughput advantage (14,000 annotations in 48 hours) without sacrificing the quality of the downstream AI.

Limitations & Future Work

  • Anatomical Complexity: The kidney is relatively easy to identify. The authors acknowledge that more challenging organs (like the pancreas, which has ill-defined borders) might still require expert oversight.
  • Pathology: This study used mostly healthy kidneys. Crowdsourcing performance on scans with large tumors or congenital deformities remains an open question.
  • Quality Control: While majority voting worked here, more sophisticated "truth discovery" algorithms (like Dawid-Skene) could further optimize the balance between the number of workers and label accuracy.

Final Thought: If we can outsource the "grunt work" of segmentation to the global crowd, we pave the way for medical AI that scales at the speed of the internet, not the speed of a clinician's schedule.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize crowdsourcing for segmenting complex and low-contrast organs like the pancreas or prostate.
  • What are the primary theoretical frameworks for "Truth Discovery" or "Label Aggregation" in medical crowdsourcing, and how do they compare to the majority voting used in this study?
  • Explore studies investigating the use of Active Learning to strategically combine expert and crowd annotations to minimize total labeling costs.
Contents
Crowdsourcing the Kidney: Scalable Medical Annotation Without the Expert Price Tag
1. TL;DR
2. Problem & Motivation: The Expert Bottleneck
3. Methodology: From Polygons to Consensus
3.1. The Workflow:
3.2. The Post-Processing Pitfall:
4. Experiments & Results: Crowd vs. Expert
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work