Optimization-Based Data Mining: Mastering the Multi-Criteria Frontier
A Family of Optimization Based Data Mining Methods
This paper provides a comprehensive survey of a family of optimization-based data mining methods centered on Multi-Criteria Programming (MCP). It introduces several evolved models including Multiple Criteria Linear Programming (MCLP) and Multiple Criteria Quadratic Programming (MCQP), achieving SOTA performance in fields like credit scoring and intrusion detection.
TL;DR
This survey explores the evolution of Multi-Criteria Programming (MCP) as a foundational tool for data mining. Unlike traditional methods that seek a single global optimum, MCP systematically addresses the "Optimization Trade-off"—such as the battle between classification accuracy (minimizing overlap) and model robustness (maximizing boundary distance). By evolving from Linear Programming to Quadratic and Fuzzy models, this family of methods provides a specialized toolkit for credit scoring, medical diagnosis, and cybersecurity.
Problem & Motivation: The Single-Criterion Trap
In the landscape of statistical learning, we are often taught to minimize a single loss function. However, the authors argue that this is an oversimplification. Real-world data mining is governed by the Bias-Variance dilemma and the Fitness-Generality trade-off.
For instance, in discriminant analysis, a "good" classifier must satisfy two conflicting goals:
- Minimize the sum of deviations (MSD): Reduce the overlap between different classes at the boundary.
- Maximize the minimum distance (MMD): Push correctly classified points as far from the boundary as possible to ensure future generalizability.
Single-criterion programming struggles to optimize these simultaneously because they are inherently contradictory. To solve this, the authors propose moving the mathematical foundation from single-objective to multi-objective optimization.
Methodology: The Evolution of MCP Models
1. Multiple Criteria Linear Programming (MCLP)
The original MCLP model (Model 1) defines two objectives: minimizing the overlapping variable and maximizing the distance variable .

To solve this efficiently, the authors utilize a compromise solution approach, transforming the bi-criteria problem into a single-objective "regret measure" model that navigates the distance towards an ideal "utopia" point .
2. Generalization to Quadratic and Multi-Group Forms
Recognizing that linear boundaries have limits, the family expanded into Multiple Criteria Quadratic Programming (MCQP). By introducing p-norms (Model 3 & 4), the researchers added stability to the classification boundary. Specifically, the inclusion of a regularization-like term (Model 5) brings it closer to the structural risk minimization seen in Support Vector Machines (SVMs), but with the added flexibility of multiple objectives.

Experiments & Results: Real-World Dominance
The survey demonstrates the power of these models across several high-impact datasets:
- Credit Card Risk: MCQP achieved 78.50%, matching or exceeding the performance of classic LDA and See5.0 Decision Trees.
- Intrusion Detection (KDDCUP-99): The Multi-Group version (MCMP) with kernels achieved a staggering 97.2% accuracy. The key benefit here was the reduction in false alarms, which is critical for network security.
- HIV-1 Neural Damage: In classifying neuronal damage, MCLP showed superior performance over Back-Propagation Neural Networks in 75% of the test cases, proving its effectiveness even in biological small-sample data.

Critical Analysis & Conclusion
Takeaway
The MCP family of methods succeeds because it explicitly acknowledges the "regret" of the decision-maker. By allowing researchers to tune the weights of MSD vs. MMD, these models offer an interpretive flexibility that black-box neural networks often lack.
Limitations & Future Work
Despite their success, two major hurdles remain:
- Computational Load: As we move toward kernelized versions of MCMP for large datasets, the optimization complexity increases significantly.
- Parametric Sensitivity: The choice of the scalar boundary is critical. The authors suggest that future work should focus on "Multi-Criteria & Multi-Constraint" (MC2) models to automatically determine optimal boundary combinations.
In conclusion, optimization-based data mining is not just a mathematical curiosity; it is a specialized framework that bridges the gap between pure statistical learning and decision-making science.
