AN EXPLAINABLE MACHINE LEARNING FRAMEWORK FOR EARLY SOFTWARE DEFECT PREDICTION: ENHANCING SOFTWARE QUALITY THROUGH INTELLIGENT CLASSIFICATION

Authors

  • Muhammad Zulqarnain Siddiqui Author
  • Hamza Bin Irfan Author
  • Muhammad Bilal Mansoor Author

Keywords:

software defect prediction; explainable artificial intelligence; XGBoost; SHAP; software metrics; class imbalance; software quality assurance

Abstract

Software defect prediction can support early allocation of verification and testing effort by identifying modules that are likely to contain defects before release. This study develops a leakage-safe explainable machine learning framework and evaluates it on four cleaned NASA Metrics Data Program benchmark datasets: CM1, JM1, KC1, and PC1. Within each training fold, the framework applies median imputation, correlation-aware feature filtering, robust scaling, and Synthetic Minority Over-sampling Technique balancing. Random Forest, radial-basis-function Support Vector Machine, XGBoost, and LightGBM are compared using stratified five-fold cross-validation and six complementary metrics. Across projects, XGBoost provided the strongest overall balance, attaining the highest macro-averaged accuracy (0.816), precision (0.435), and F1-score (0.402). Random Forest achieved the highest ROC-AUC (0.761), LightGBM achieved the highest precision-recall AUC (0.429), and Support Vector Machine achieved the highest recall (0.470), demonstrating that the preferred model depends on the operational cost of missed defects and false alarms. A Friedman test detected an overall model difference for ROC-AUC (chi-square = 8.40, p = 0.038), whereas differences for the remaining metrics were not statistically significant across the four projects. SHAP analysis of XGBoost on JM1 identified blank lines, numbers of unique operators and operands, design complexity, Halstead difficulty, cyclomatic complexity, and total lines of code as influential predictors. The framework therefore combines competitive prediction with transparent global and module-level explanations, while avoiding optimistic leakage from preprocessing or class balancing. The results support explainable defect prediction as a decision-support mechanism for risk-based testing rather than as an automatic replacement for engineering judgment.

Downloads

Published

2026-05-31