Stratification in SDP: Why the Interaction of Under- and Over-sampling is the Key to Reliability

10659_Evaluating Stratification Alternatives to Improve Software Defect Prediction.

Summary
Problem
Method
Results
Takeaways
Abstract

The paper evaluates the impact of stratification techniques (under-sampling and over-sampling) on Software Defect Prediction (SDP) using a factorial ANOVA design. It demonstrates that while over-sampling alone isn't statistically significant, under-sampling and the interaction between under- and over-sampling significantly improve classification accuracy (measured by Cohen’s Kappa) across multiple NASA and open-source datasets.

TL;DR

Predicting software defects is a classic "needle in a haystack" problem. This paper provides a rigorous statistical foundation proving that we shouldn't just over-sample or under-sample—we must do both. By applying a factorial ANOVA design to seven large-scale datasets, the authors reveal that the interaction between these two methods is where the real performance gains reside.

Background: The Imbalanced Reality of Software Metrics

Software Defect Prediction (SDP) aims to identify modules likely to fail based on metrics like Code Complexity or Inheritance Depth. However, most modules in a healthy system are not defective. This creates a "skewed" dataset where a classifier can achieve 90% accuracy simply by predicting "no defect" every time—a result that is 0% useful for quality assurance.

The "Either/Or" Debate

In the machine learning community, a debate has raged:

  • Under-sampling advocates argue it reduces noise and speeds up training.
  • Over-sampling (SMOTE) advocates argue it preserves valuable information by creating synthetic "faulty" examples.

This paper moves beyond the debate to ask: Why not both?

Methodology: High-Stakes Statistics

The authors employed a Blocked Full-Factorial Design. They didn't just test methods; they tested the levels of intensity for each method across multiple NASA datasets (KC1, JM1, etc.) and the Mozilla open-source project.

The Core Framework

  1. Metric Selection: Utilized CK metrics (OO-specific) and procedural metrics (McCabe complexity).
  2. Algorithm: Used the C4.5 Decision Tree as a standardized baseline.
  3. Factorial Design:
    • Factor A (Under-sampling): 0% to 80% removal.
    • Factor B (Over-sampling): 0% to 400% synthetic addition using SMOTE and SMOTE-SN.
  4. Metric of Success: Cohen's Kappa (), which accounts for agreement occurring by chance.

Analysis of Variance (ANOVA) Design

Key Insights from the ANOVA

The statistical results were surprising. In the ANOVA table above, we see:

  • Dataset (Blocking Factor): Highly significant, confirming that results are heavily project-dependent.
  • Under-sampling: Significantly improved models on its own.
  • Over-sampling: Statistically insignificant as a main effect!
  • Interaction (Under * Over): Highly significant.

This suggests that SMOTE (over-sampling) is not a "magic bullet" on its own, but when the majority class is simultaneously pruned (under-sampling), SMOTE's synthetic examples become much more effective at defining clear decision boundaries.

Experimental Showdown: SMOTE vs. SMOTE-SN

The study also tested a newer variant, SMOTE-SN (Surrounding Neighbors), which uses symmetry and proximity to generate samples. While SMOTE-SN showed potential, the underlying lesson remained the same: the synergy between reducing the "normal" class and augmenting the "defect" class is vital.

Tukey's HSD Test Results

Critical Analysis & Conclusion

Takeaway

If you are building a software reliability model, hybridization is mandatory. Relying exclusively on SMOTE may provide no significant benefit over a baseline, while a combined approach systematically improves the statistic.

Limitations

  • Algorithmic Bias: The study exclusively uses C4.5. While C4.5 is a robust benchmark, modern ensembles like XGBoost or Random Forests might react differently to stratification.
  • Dataset Source: Six out of seven datasets came from NASA. Despite the inclusion of Mozilla, further cross-industry verification is needed.

Future Outlook

The authors suggest moving toward Response Surface Methodology (RSM). Instead of choosing arbitrary levels (e.g., 20% under-sampling), future tools could automatically "surface" the mathematically optimal balance of sampling for a specific codebase's unique skew.

Find Similar Papers

Try Our Examples

  • Search for recent studies after 2012 that use Response Surface Methodology (RSM) to optimize the specific percentages of under-sampling and over-sampling in software defect prediction.
  • Which paper first introduced the SMOTE-SN (Surrounding Neighbors) variation, and how does its synthetic instance generation logic differ from the original SMOTE algorithm?
  • Explore comparative research that evaluates cost-sensitive learning versus combined resampling (hybrid stratification) in high-skew domains like cybersecurity or medical diagnosis.
Contents
Stratification in SDP: Why the Interaction of Under- and Over-sampling is the Key to Reliability
1. TL;DR
2. Background: The Imbalanced Reality of Software Metrics
3. The "Either/Or" Debate
4. Methodology: High-Stakes Statistics
4.1. The Core Framework
5. Key Insights from the ANOVA
6. Experimental Showdown: SMOTE vs. SMOTE-SN
7. Critical Analysis & Conclusion
7.1. Takeaway
7.2. Limitations
7.3. Future Outlook