FH-GBML-CS: Optimizing Genetic Fuzzy Systems for High Imbalance through Cost-Sensitivity
A first approach for cost-sensitive classification with linguistic Genetic Fuzzy Systems in imbalanced data-sets
The paper introduces FH-GBML-CS, a cost-sensitive genetic fuzzy system designed for high imbalanced classification. By integrating misclassification costs directly into the fitness function and rule weight calculation, it achieves SOTA performance comparable to SMOTE-based preprocessing.
TL;DR
Addressing class imbalance usually involves "rebalancing" data (e.g., SMOTE), but this often leads to bloated models and long training cycles. This paper proposes FH-GBML-CS, an algorithmic-level solution that embeds misclassification costs directly into a Genetic Fuzzy System. The result? A model that matches the accuracy of SMOTE but is 3x faster to train and 5x more interpretable.
Background: The Price of Imbalance
In many real-world scenarios—from parasite detection in medicine to mine detection in radar—the "minority class" is the one we care about most. However, standard Machine Learning algorithms are "lazy" optimizers; they maximize global accuracy, which often means ignoring the minority class entirely to achieve 99% accuracy by simply labeling everything as "majority."
While Data-level solutions (Sampling) and Algorithmic-level solutions exist, this paper argues for a Cost-Sensitive approach. By assigning a higher penalty to missing a minority instance, we can guide the learning process to prioritize what's actually important.
Methodology: Engineering Cost-Sensitivity into Fuzzy Rules
The authors chose the Fuzzy Hybrid Genetics-Based Machine Learning (FH-GBML) algorithm as their foundation. They introduced three surgical modifications to make it cost-aware:
- Cost-Driven Fitness: Instead of using the count of correctly classified instances, the evolutionary "fitness" of a rule set is determined by minimizing a cost metric: .
- CS-PCF (Cost-Sensitive Penalized Certainty Factor): Rule weights are no longer just about frequency. The new heuristic incorporates the misclassification cost () of each instance, ensuring that rules favoring the minority class carry more weight.
- Label Selection: When assigning a class to a fuzzy rule, the system picks the class that maximizes the product of , shifting the decision boundary in favor of the minority class.
The linguistic rule structure used in the proposed FRBCS system.
Experiments & Results
The researchers tested the approach on 22 "High Imbalance" datasets (where the Imbalance Ratio > 9).
1. Performance (AUC)
The Area Under the ROC Curve (AUC) is the gold standard for imbalanced data. FH-GBML-CS achieved 82.35%, a massive leap from the standard version's 58.92%.
2. The Efficiency Factor
One of the most striking findings is the efficiency of the cost-sensitive approach compared to SMOTE (Synthetic Minority Over-sampling Technique).
- FH-GBML+SMOTE: 1140 seconds.
- FH-GBML-CS: 342 seconds.
By avoiding the generation of synthetic data, the algorithm maintains a smaller, more manageable search space for the Genetic Algorithm.
Summary of results showing Average AUC, Rule count, and Run Time.
3. Interpretability (Rule Complexity)
In a "Linguistic" Fuzzy System, the goal is often human readability. FH-GBML-CS produced models with an average of only 6.89 rules, compared to 33.95 for the baseline. This degree of compression makes the model significantly easier for domain experts to audit.
Critical Analysis & Conclusion
Takeaway
FH-GBML-CS proves that we don't always need more data; we need better-targeted optimization. By moving the "pressure" of the imbalance into the algorithm's internal scoring and weighting mechanisms, the authors created a leaner, faster, and more accurate classifier.
Limitations & Future Work
The study primarily focuses on binary classification. As the authors suggest, the next frontier is extending this cost-sensitive logic to multi-class imbalances, where the cost matrix becomes significantly more complex. Furthermore, testing these "lightweight" rule sets against modern Deep Learning benchmarks for tabular data could further validate the "Interpretability vs. Accuracy" trade-off.
Academic Note: This work highlights the enduring relevance of Genetic Fuzzy Systems in niche domains where transparency and high-stakes decision-making are paramount.
