CBEUS: Solving the Imbalance Bottleneck in Bankruptcy Prediction via Evolutionary Intelligence
Expert Systems With Applications
This paper introduces a Cluster-Based Evolutionary Undersampling (CBEUS) approach for bankruptcy prediction. It combines K-means clustering with Genetic Algorithms (GA) to optimize Artificial Neural Network (ANN) training sets by selectively removing majority-class instances.
TL;DR
Corporate bankruptcy prediction is a high-stakes task often crippled by "Data Imbalance"—where healthy firms vastly outnumber failing ones. This paper presents CBEUS (Cluster-Based Evolutionary Undersampling), a hybrid model that uses K-means clustering and Genetic Algorithms to surgically prune the majority class. By optimizing instance selection and neural network weights simultaneously, the authors achieved a massive leap in G-Mean (from 17.21% to 84.26%), ensuring that minority bankruptcy cases are no longer "ignored" by the AI.
The Imbalance Pain Point: Why Accuracy is a Lie
In financial datasets, bankruptcy is a rare event. If 95% of firms are healthy, a model can achieve 95% accuracy by simply predicting "No Bankruptcy" for everyone. This is the Accuracy Paradox.
- Prior Work Failures: Random undersampling (RUS) loses valuable data, while oversampling (SMOTE) can introduce artificial noise.
- The Insight: Not all non-bankrupt firms are equally useful for training. Some are "typical," while others are "noisy" or "outliers" that confuse the decision boundary. The authors propose that we should categorize the majority class first, then use evolution to find the best boundary for each sub-group.
Methodology: The GA-ANN Hybrid Architecture
The proposed CBEUS framework operates in three distinct phases:
- Structural Recognition (Clustering): The non-bankrupt firms are grouped using K-means. The "Silhouette statistic" is used to determine the optimal number of clusters (found to be ).
- Evolutionary Pruning: A Genetic Algorithm (GA) searches for specific distance thresholds () for each cluster. If a firm's distance from its cluster centroid exceeds the threshold, it is labeled as "noise" and removed.
- Simultaneous Optimization: Unlike traditional methods that treat sampling and modeling separately, the GA here optimizes both the selection rules and the ANN connection weights at the same time.

The Chromosome Structure
The GA encodes a complex search space: Where represents the cluster thresholds and represents the neural network weights. By using the G-Mean as the fitness function, the model is forced to maximize the balance between Sensitivity (identifying failed firms) and Specificity (identifying healthy firms).
Experiments & SOTA Results
The researchers tested the model on 22,500 Korean manufacturing firms. The bankruptcy rate was a mere 5.9%, representing a significant "extreme imbalance" challenge.
Performance Comparison
The CBEUS method was compared against standard ANNs, Random Undersampling (RUS), and standard Evolutionary Undersampling (EUS).
| Metric | ANN (None) | ANN (RUS) | GA-ANN (CBEUS) |
|---|---|---|---|
| Sensitivity | 2.96% | 69.63% | 87.41% |
| G-Mean | 17.21% | 78.99% | 84.26% |
| H-Measure | 42.59 | 47.60 | 56.16 |

The results show that CBEUS doesn't just improve accuracy—it fundamentally shifts the model's ability to "see" the minority class. The use of the H-measure further validates that the model is robust against varying misclassification costs.
Critical Analysis & Takeaways
The brilliance of this work lies in its Rule-Format Representation. Instead of a black-box sampler, the model generates an interpretable rule:
IF [Distance from Cluster_1 < 0.501] AND [Distance from Cluster_2 < 0.520] ... THEN Select Instance.
Limitations:
- Computational Cost: GA is notoriously slow compared to gradient-based methods. Searching for both thresholds and weights simultaneously in a massive big data environment could lead to the "curse of dimensionality."
- Clustering Sensitivity: The model's success is heavily reliant on the initial K-means quality.
Future Outlook: This research paves the way for "Self-Organizing" datasets where the AI actively participates in its own data cleaning. For financial institutions, this means more reliable early-warning systems and significantly lower risks of undetected corporate failures.
