Accelerating Medical Diagnosis: A GPU-Based BP Neural Network for Healthcare Big Data
A GPU-Based Training of BP Neural Network for Healthcare Data Analysis
This paper introduces a GPU-accelerated Back Propagation (BP) neural network designed for high-efficiency healthcare data analysis, specifically targeting tumor diagnosis. By parallelizing the weight updates and neuron computations, the system achieves significant speedups compared to traditional CPU-based implementations.
TL;DR
In the era of medical big data, processing speed is as critical as diagnostic accuracy. This paper presents a GPU-parallelized Back Propagation (BP) neural network that overcomes the latency bottlenecks of traditional CPU training. By leveraging the massively parallel architecture of GPUs, the authors transform tumor diagnosis—specifically breast cancer detection—into a real-time process, achieving a high throughput of nearly 125 FPS.
Problem & Motivation: The Latency Gap in Healthcare
While machine learning has become a cornerstone of pathological diagnosis, the "Big Data" nature of modern medical records poses a significant challenge. Traditional CPU execution of training algorithms is inherently sequential, leading to:
- Inability to Update: Models are often trained once and left static because retraining is too costly.
- Latency: Analyzing large-scale historical records for real-time clinical support becomes impossible.
- Poor Adaptability: Static models cannot easily integrate new patient data to refine their predictive capabilities.
The authors' insight is straightforward yet powerful: the mathematical operations of a neural network—dot products, activations, and gradient updates—are perfectly suited for the SIMD (Single Instruction, Multiple Data) architecture of GPUs.
Methodology: Parallelizing the Three-Layer Perceptron
The proposed system utilizes a three-layer BP neural network (Input, Hidden, Output).
1. The Architecture
The model is specifically tuned for breast cancer datasets:
- Input Layer: 9 neurons (representing features like bare nuclei, bland chromatin, and mitoses).
- Hidden Layer: 5 neurons utilizing the Sigmoid activation function.
- Output Layer: 2 neurons (classifying as malignant or benign).
2. GPU Parallel Strategy
Instead of calculating each neuron's output one by one, the system assigns computations to distinct GPU threads.
- Forward Propagation: The Sigmoid function is executed in parallel across the GPU grid.
- Weight Updates: The loss function and the subsequent gradient descent are computed such that all parameters are updated simultaneously in each iteration.
Figure 1: The structure of the multi-layer feed-forward network used for tumor classification.
Experiments & Results
The system was tested using the Wisconsin Breast Cancer dataset from the UCI repository.
Performance Benchmarks
The most striking result is the efficiency gain. The GPU-based model reached 124.99 FPS during the training phase. This speed allows for iterative retraining as new data points are introduced, a feat difficult for traditional CPU setups on the same hardware class (Core i5).
Accuracy and Convergence
The experiment evaluated the model across different learning rates. The results demonstrate that the model converges reliably; however, the choice of learning rate remains critical to avoid local minima or slow convergence in the parallel environment.
Figure 2: Accuracy tracking across different learning rates and iteration numbers, demonstrating stable convergence.
Critical Analysis & Conclusion
Takeaway
This work serves as a foundational bridge between classical machine learning and high-performance computing (HPC) in medicine. By moving beyond CPU-bound training, it enables Synchronous Training—where models learn while they are being used for testing and diagnosis.
Limitations & Future Work
- Architectural Simplicity: While effective for tabular data (9 attributes), modern deep learning handles much larger feature sets (images/genomics). The next step would be scaling this parallel approach to Deep CNNs or Transformers for medical imaging.
- Hardware Dependency: The efficiency is tied to the memory bandwidth between the CPU and GPU. Future optimizations could include minimizing data transfer overhead (PCIe bottlenecks).
In conclusion, the integration of GPU programming into BP networks is not just a speed upgrade; it is a necessary evolution for making AI-driven healthcare responsive and adaptable at scale.
