CRDT: Overcoming Information Gain Bias in Healthcare Data Mining
CRDT: Correlation Ratio Based Decision Tree Model for Healthcare Data Mining
The paper introduces CRDT (Correlation Ratio Based Decision Tree), a novel classification model tailored for healthcare data mining. It replaces the traditional Information Gain (IG) metric with Correlation Ratio (CR) to achieve SOTA-comparable performance while eliminating the inherent bias toward attributes with numerous distinct values.
TL;DR
The Correlation Ratio Based Decision Tree (CRDT) addresses a fundamental flaw in traditional ID3/C4.5 algorithms: the bias toward high-cardinality attributes. By utilizing a statistical Correlation Ratio (CR) for splitting, this model provides more robust classifications for healthcare datasets like Hepatitis and Liver disease, where feature significance is often overshadowed by the number of distinct values.
The "Cardinality Trap" in Medical Data
Information Gain (IG) is the bedrock of many Decision Tree models. However, it possesses an "Achilles' heel"—it mathematically prefers attributes with more distinct values. In a clinical setting, an attribute like "Patient ID" would yield the highest Information Gain because it creates perfectly pure (but useless) partitions.
Authors Roy et al. argue that healthcare data is uniquely susceptible to this. Medical records often mix binary flags (smoker/non-smoker) with high-variance categorical data (blood types, symptoms). IG-based trees frequently pick the "noisier" high-cardinality features, leading to models that fail to generalize on new patients.
Methodology: The Power of Correlation Ratio
The core innovation lies in the transition from Entropy to Correlation Ratio (CR).
1. Intuition Behind CR
Unlike IG, which looks at the reduction in uncertainty, CR looks at the dispersion of values. A significant attribute is one where the average value within a specific outcome class (e.g., "Hepatitis Positive") is remarkably different from the overall population average.
2. The Logic
The algorithm follows a classic recursive partitioning structure but calculates the CR for each attribute at every node.
- Numerator: Dispersion among individual classes.
- Denominator: Dispersion across the whole population.
The model chooses the attribute that maximizes this ratio, ensuring that the split is based on true class-relevance rather than just the "splitting power" of many distinct values.
(Note: Refer to Algorithm 1 and 2 in the paper for the recursive construction logic and the adaptation of CR for nominal attributes.)
Experimental Results: Where CRDT Shines
The researchers tested CRDT against IG across eight benchmark datasets from the UCI repository.
Key Findings:
- Superiority in Small/Complex Datasets: CRDT outperformed IG in the Hepatitis dataset (73.78% vs 71.19%) and the Indian Liver Patient Dataset (ILPD). These datasets are characterized by features with differing numbers of distinct values.
- Consistency: In datasets where attributes had uniform cardinalities (like Spect-heart), CRDT matched IG's performance (74.33%) exactly, proving it is a safe "drop-in" replacement.
- Robustness: The results indicate that as the dataset becomes "messier" with varying attribute types, CRDT’s lack of bias becomes a significant advantage.
(Note: Table V in the paper details the cross-validation results across all 8 healthcare benchmarks.)
Critical Insight: When to Use CRDT?
The paper suggests a "complementary" relationship. CRDT is not a "silver bullet" to replace all Decision Trees, but it is a superior choice when:
- The dataset contains a mix of categorical and discretized numerical data.
- Prior IG models show signs of overfitting on high-cardinality features.
- The sample size is relatively small (like the Hepatitis or Statlog datasets), where every split decision critically impacts the final accuracy.
Conclusion and Future Directions
The CRDT model provides a mathematically grounded solution to the bias inherent in Information Gain. By shifting the focus to Correlation Ratios, the authors have created a tool that respects the biological significance of medical features over their statistical distribution. Future work could involve scaling this to ensemble methods like Random Forests to see if the "CR-Forest" outperforms the standard IG-based versions.
Author Affiliations: Smita Roy (Central University of Bihar), Samrat Mondal & Asif Ekbal (IIT Patna).
