IDTC: Enhancing Heart Disease Diagnosis through Multivariate Decision Trees
A Knowledgeable Decision Tree Classification Model for Multivariate Heart Disease Data-A Boon to Healthcare
The paper introduces the Intelligent Decision Tree Construction (IDTC) algorithm, a novel classification model designed to handle multivariate Coronary Artery Disease (CAD) datasets. By extending traditional decision tree mechanics to manage complex attribute interdependencies, the method achieves improved diagnostic performance on the UCI Cleveland dataset.
TL;DR
This research presents the Intelligent Decision Tree Construction (IDTC) algorithm, specifically engineered for the complexities of multivariate heart disease data. By optimizing split-point selection and employing rigorous pruning, the model surpasses traditional C4.5-based classifiers in accuracy, sensitivity, and specificity on the standard UCI Cleveland heart disease dataset.
Problem & Motivation: The Chaos of Multivariate Clinical Data
In a smart hospital environment, medical diagnosis relies heavily on multivariate clinical data—sets where multiple variables are tracked simultaneously. However, classifying this data is notoriously difficult due to:
- The Boundary Effect: Mutual dependence between attributes distorts the data space, making near points appear far and causing significant errors in distance-based classification.
- Subjectivity: Older expert systems require manual knowledge entry from physicians, which is time-consuming and prone to human bias.
- Data Impurity: Real-world medical data is often "impure," containing noise, errors, and missing values that standard algorithms handle poorly.
The authors' insight was to move beyond the statistical simplicity of the C4.5 algorithm and create a system that calculates split points () based on the inherent stability of the multivariate system.
Methodology: The IDTC Algorithm
The core of the proposed work is the IDTC Algorithm. Unlike traditional trees that partition data based on simple entropy or information gain for single variables, IDTC treats the dataset as a multivariate system.
Key Execution Steps:
- Attribute Assignment: Attributes are mapped as a set .
- Split Point Calculation: The algorithm identifies a point value () for each attribute. The split point is chosen empirically to represent the "tangent to the stable multivariate decision system."
- Recursive Partitioning: Tuples are moved to "Left" or "Right" branches based on an input value comparison ().
- Automatic Pruning: To prevent overfitting (a common issue with deep trees), the algorithm is pruned up to 7 levels.
Fig 1. Logical flow of the proposed IDTC classification model.
Experiments and Performance
The model was validated using the UCI Cleveland Heart Disease dataset, which contains 303 records and 13 critical attributes (such as age, sex, chest pain type, and blood pressure).
SOTA Comparison
The IDTC model was compared against the Traditional Decision Tree classifier. The results across various metrics showed a consistent lead for the Multivariate approach:
- Sensitivity (Recall): Achieved 73.1%, significantly higher than the traditional 69.1%, meaning fewer sick patients were missed.
- Specificity: Reached 82.3%, improving the ability to correctly identify healthy individuals.
- Overall Accuracy: Stabilized at 76.78%, providing a more reliable overall diagnostic tool.
Fig 2. Visualization of the tree construction in the MATLAB environment.
Deep Insight & Conclusion
The success of the IDTC model lies in its recognition that heart disease diagnosis is not a linear set of independent factors but a complex web of interdependent biological signals. By using a split-point strategy that considers the "stability" of the multivariate system, the model effectively mitigates the "boundary effect" that hampers traditional distance-based methods.
Limitations & Future Work
- Computational Complexity: Calculating multivariate split points is more taxing than univariate splits.
- Data Scale: While effective on the Cleveland dataset, the algorithm needs testing on larger, more diverse "big data" healthcare sets.
- Future Path: The authors intend to integrate more advanced Machine Learning techniques to further refine the split-point logic and explore even higher-dimensional data.
In summary, IDTC offers a robust, interpretable "boon to healthcare" by automating the extraction of reliable diagnostic rules from complex raw medical data.
