Stratification in SDP: Why the Interaction of Under- and Over-sampling is the Key to Reliability
10659_Evaluating Stratification Alternatives to Improve Software Defect Prediction.
The paper evaluates the impact of stratification techniques (under-sampling and over-sampling) on Software Defect Prediction (SDP) using a factorial ANOVA design. It demonstrates that while over-sampling alone isn't statistically significant, under-sampling and the interaction between under- and over-sampling significantly improve classification accuracy (measured by Cohen’s Kappa) across multiple NASA and open-source datasets.
TL;DR
Predicting software defects is a classic "needle in a haystack" problem. This paper provides a rigorous statistical foundation proving that we shouldn't just over-sample or under-sample—we must do both. By applying a factorial ANOVA design to seven large-scale datasets, the authors reveal that the interaction between these two methods is where the real performance gains reside.
Background: The Imbalanced Reality of Software Metrics
Software Defect Prediction (SDP) aims to identify modules likely to fail based on metrics like Code Complexity or Inheritance Depth. However, most modules in a healthy system are not defective. This creates a "skewed" dataset where a classifier can achieve 90% accuracy simply by predicting "no defect" every time—a result that is 0% useful for quality assurance.
The "Either/Or" Debate
In the machine learning community, a debate has raged:
- Under-sampling advocates argue it reduces noise and speeds up training.
- Over-sampling (SMOTE) advocates argue it preserves valuable information by creating synthetic "faulty" examples.
This paper moves beyond the debate to ask: Why not both?
Methodology: High-Stakes Statistics
The authors employed a Blocked Full-Factorial Design. They didn't just test methods; they tested the levels of intensity for each method across multiple NASA datasets (KC1, JM1, etc.) and the Mozilla open-source project.
The Core Framework
- Metric Selection: Utilized CK metrics (OO-specific) and procedural metrics (McCabe complexity).
- Algorithm: Used the C4.5 Decision Tree as a standardized baseline.
- Factorial Design:
- Factor A (Under-sampling): 0% to 80% removal.
- Factor B (Over-sampling): 0% to 400% synthetic addition using SMOTE and SMOTE-SN.
- Metric of Success: Cohen's Kappa (), which accounts for agreement occurring by chance.

Key Insights from the ANOVA
The statistical results were surprising. In the ANOVA table above, we see:
- Dataset (Blocking Factor): Highly significant, confirming that results are heavily project-dependent.
- Under-sampling: Significantly improved models on its own.
- Over-sampling: Statistically insignificant as a main effect!
- Interaction (Under * Over): Highly significant.
This suggests that SMOTE (over-sampling) is not a "magic bullet" on its own, but when the majority class is simultaneously pruned (under-sampling), SMOTE's synthetic examples become much more effective at defining clear decision boundaries.
Experimental Showdown: SMOTE vs. SMOTE-SN
The study also tested a newer variant, SMOTE-SN (Surrounding Neighbors), which uses symmetry and proximity to generate samples. While SMOTE-SN showed potential, the underlying lesson remained the same: the synergy between reducing the "normal" class and augmenting the "defect" class is vital.

Critical Analysis & Conclusion
Takeaway
If you are building a software reliability model, hybridization is mandatory. Relying exclusively on SMOTE may provide no significant benefit over a baseline, while a combined approach systematically improves the statistic.
Limitations
- Algorithmic Bias: The study exclusively uses C4.5. While C4.5 is a robust benchmark, modern ensembles like XGBoost or Random Forests might react differently to stratification.
- Dataset Source: Six out of seven datasets came from NASA. Despite the inclusion of Mozilla, further cross-industry verification is needed.
Future Outlook
The authors suggest moving toward Response Surface Methodology (RSM). Instead of choosing arbitrary levels (e.g., 20% under-sampling), future tools could automatically "surface" the mathematically optimal balance of sampling for a specific codebase's unique skew.
