EVALUATING EXPLAINABLE ARTIFICIAL INTELLIGENCE FOR HIGH-STAKES DECISION MAKING: AN EMPIRICAL ASSESSMENT OF PREDICTIVE ACCURACY, ALGORITHMIC FAIRNESS, AND MODEL INTERPRETABILITY

Authors

  • Saba Jabbar Author
  • Gulzar Ali Brohi Author
  • Muhammad Mudaser Author
  • Atif Saeed Author
  • Nazia Azim Author

Keywords:

explainable artificial intelligence; XAI; algorithmic fairness; interpretability; COMPAS; SHAP; calibration; high-stakes decision making; responsible AI

Abstract

High-stakes artificial intelligence requires evaluation beyond predictive accuracy because errors, calibration, demographic disparities, and explanation quality can have consequential effects. This study evaluated logistic regression, a shallow decision tree, random forest, XGBoost, and an interpretable spline generalized additive model (GAM) using the ProPublica COMPAS two-year recidivism dataset (N=7,214). A leakage-conscious feature set included age, sex, prior-offense count, juvenile-offense counts, and charge degree, while race was reserved for fairness auditing. The primary 75:25 holdout was supplemented by repeated stratified five-fold cross-validation with three repeats. Mean cross-validated ROC-AUC was 0.724 for logistic regression, 0.718 for the decision tree, 0.729 for random forest, and approximately 0.731 for both XGBoost and the spline GAM. Calibration was broadly comparable (held-out ECE 0.020-0.037), and threshold analysis from 0.20 to 0.80 showed that measured racial disparities changed materially with the operating point. At threshold 0.50, all models produced higher positive-prediction and false-positive rates for African-American than Caucasian observations. A secondary sex-based analysis also identified non-trivial disparities. XGBoost explanations were dominated by prior-offense count and age; global SHAP rankings were highly stable across bootstrap refits (mean pairwise Spearman rho=0.973; mean top-five overlap=1.00). These findings show that model complexity produced limited discrimination gains and did not eliminate demographic disparities. High-stakes model evaluation should therefore jointly examine discrimination, calibration, subgroup error burdens, threshold sensitivity, and explanation stability rather than relying on accuracy or explanation display alone.

Downloads

Published

2026-09-19