Accelerating Medical Diagnosis: A GPU-Based BP Neural Network for Healthcare Big Data

A GPU-Based Training of BP Neural Network for Healthcare Data Analysis

2018-11-28
Wei Song, Shuanghui Zou, Yifei Tian, Simon Fong
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a GPU-accelerated Back Propagation (BP) neural network designed for high-efficiency healthcare data analysis, specifically targeting tumor diagnosis. By parallelizing the weight updates and neuron computations, the system achieves significant speedups compared to traditional CPU-based implementations.

TL;DR

In the era of medical big data, processing speed is as critical as diagnostic accuracy. This paper presents a GPU-parallelized Back Propagation (BP) neural network that overcomes the latency bottlenecks of traditional CPU training. By leveraging the massively parallel architecture of GPUs, the authors transform tumor diagnosis—specifically breast cancer detection—into a real-time process, achieving a high throughput of nearly 125 FPS.

Problem & Motivation: The Latency Gap in Healthcare

While machine learning has become a cornerstone of pathological diagnosis, the "Big Data" nature of modern medical records poses a significant challenge. Traditional CPU execution of training algorithms is inherently sequential, leading to:

  • Inability to Update: Models are often trained once and left static because retraining is too costly.
  • Latency: Analyzing large-scale historical records for real-time clinical support becomes impossible.
  • Poor Adaptability: Static models cannot easily integrate new patient data to refine their predictive capabilities.

The authors' insight is straightforward yet powerful: the mathematical operations of a neural network—dot products, activations, and gradient updates—are perfectly suited for the SIMD (Single Instruction, Multiple Data) architecture of GPUs.

Methodology: Parallelizing the Three-Layer Perceptron

The proposed system utilizes a three-layer BP neural network (Input, Hidden, Output).

1. The Architecture

The model is specifically tuned for breast cancer datasets:

  • Input Layer: 9 neurons (representing features like bare nuclei, bland chromatin, and mitoses).
  • Hidden Layer: 5 neurons utilizing the Sigmoid activation function.
  • Output Layer: 2 neurons (classifying as malignant or benign).

2. GPU Parallel Strategy

Instead of calculating each neuron's output one by one, the system assigns computations to distinct GPU threads.

  • Forward Propagation: The Sigmoid function is executed in parallel across the GPU grid.
  • Weight Updates: The loss function and the subsequent gradient descent are computed such that all parameters are updated simultaneously in each iteration.

Model Architecture Figure 1: The structure of the multi-layer feed-forward network used for tumor classification.

Experiments & Results

The system was tested using the Wisconsin Breast Cancer dataset from the UCI repository.

Performance Benchmarks

The most striking result is the efficiency gain. The GPU-based model reached 124.99 FPS during the training phase. This speed allows for iterative retraining as new data points are introduced, a feat difficult for traditional CPU setups on the same hardware class (Core i5).

Accuracy and Convergence

The experiment evaluated the model across different learning rates. The results demonstrate that the model converges reliably; however, the choice of learning rate remains critical to avoid local minima or slow convergence in the parallel environment.

Experimental Results Figure 2: Accuracy tracking across different learning rates and iteration numbers, demonstrating stable convergence.

Critical Analysis & Conclusion

Takeaway

This work serves as a foundational bridge between classical machine learning and high-performance computing (HPC) in medicine. By moving beyond CPU-bound training, it enables Synchronous Training—where models learn while they are being used for testing and diagnosis.

Limitations & Future Work

  • Architectural Simplicity: While effective for tabular data (9 attributes), modern deep learning handles much larger feature sets (images/genomics). The next step would be scaling this parallel approach to Deep CNNs or Transformers for medical imaging.
  • Hardware Dependency: The efficiency is tied to the memory bandwidth between the CPU and GPU. Future optimizations could include minimizing data transfer overhead (PCIe bottlenecks).

In conclusion, the integration of GPU programming into BP networks is not just a speed upgrade; it is a necessary evolution for making AI-driven healthcare responsive and adaptable at scale.

Find Similar Papers

Try Our Examples

  • Find recent papers that compare GPU-based parallelization of classic neural networks with modern GPGPU frameworks like CUDA or ROCm in medical informatics.
  • What is the definitive paper on the parallelization of the Sigmoid activation function in early GPU computing, and how does this paper's implementation differ?
  • Have there been studies applying this parallel BP neural network approach to multi-modal healthcare data, such as combining tabular electronic health records with medical imaging?
Contents
Accelerating Medical Diagnosis: A GPU-Based BP Neural Network for Healthcare Big Data
1. TL;DR
2. Problem & Motivation: The Latency Gap in Healthcare
3. Methodology: Parallelizing the Three-Layer Perceptron
3.1. 1. The Architecture
3.2. 2. GPU Parallel Strategy
4. Experiments & Results
4.1. Performance Benchmarks
4.2. Accuracy and Convergence
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work