EconPapers    
Economics at your fingertips  
 

Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support

Oluwaseun Adebayo Bamodu, Sumaiya Nezam and Chen-Chih Chung

PLOS Digital Health, 2026, vol. 5, issue 9, 1-17

Abstract: Breast cancer remains the most commonly diagnosed malignancy among women globally, with disproportionately higher mortality rates in low- and middle-income countries (LMICs) where diagnostic delays and limited specialist pathology capacity are widespread. While machine learning (ML) approaches achieve strong predictive performance for cancer classification, algorithmic opacity and absence of interpretability frameworks tailored to resource-constrained environments have impeded clinical adoption. This study bridges the translational gap between predictive accuracy and clinical utility by developing an explainable artificial intelligence (XAI) framework specifically designed for breast cancer diagnosis in underserved healthcare settings. Using the Wisconsin Breast Cancer Diagnostic Dataset (569 fine-needle aspirate cytological specimens with 30 nuclear morphometric features), we systematically benchmarked eight supervised classification algorithms: Logistic Regression, Random Forest, XGBoost, LightGBM, Support Vector Machine (SVM), Gradient Boosting, Decision Tree, and K-Nearest Neighbors, using stratified 10-fold cross-validation and an independent hold-out test set (80:20 split). Performance was evaluated across discriminative and probabilistic metrics, including AUC-ROC, F1-score, Matthews Correlation Coefficient (MCC), and Brier score, and interpretability was operationalized through SHapley Additive exPlanations (SHAP) analysis with global feature importance, cross-model consensus ranking, and individual-level dependence characterization. All ensemble and regularized models achieved test-set AUCs above 0.98, with XGBoost and SVM attaining the highest AUC of 0.996, and Logistic Regression the highest accuracy (98.25%) and MCC (0.962). SHAP analysis consistently identified worst perimeter, worst concave points, and worst area as the dominant predictors, with strong concordance across gradient-boosted models (pairwise Spearman rho: XGBoost–LightGBM 0.86, XGBoost–Random Forest 0.82, Random Forest–LightGBM 0.67). Logistic Regression also demonstrated superior probability calibration, a critical requirement for clinical risk stratification. Collectively, these findings deliver a reproducible, transparent framework whose SHAP-derived signatures align with established cytopathological principles, supporting responsible integration of interpretable ML into resource-limited diagnostic workflows and providing a template for equitable AI deployment in global oncology.Author summary: Breast cancer is the most commonly diagnosed cancer among women worldwide, yet survival rates remain dramatically lower in low- and middle-income countries (LMICs) compared to high-income nations, largely due to limited access to specialist diagnostic expertise. Machine learning (ML) holds promise for supporting cancer diagnosis in these settings, but most high-performing ML systems function as difficult-to-interpret ‘black boxes,’ undermining clinician trust and adoption, particularly in environments where algorithmic outputs cannot be readily verified by specialist pathologists. In this study, we developed and evaluated an interpretable ML framework for breast cancer classification using fine-needle aspiration cytology data routinely collected in resource-limited settings. By systematically comparing eight ML algorithms and applying SHAP (SHapley Additive exPlanations)-based explainability analysis, we identified which cellular features most strongly influence diagnostic predictions and demonstrated that these features align with established cytopathological criteria. Crucially, we show that simpler, more interpretable models achieve performance comparable to complex ensembles while offering superior probability calibration, a critical property for clinical triage decisions. Our findings support the use of transparent, computationally lightweight ML systems as viable decision-support tools in underserved healthcare environments, advancing the goal of equitable cancer diagnosis globally.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001706 (text/html)
https://journals.plos.org/digitalhealth/article/fi ... 01706&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pdig00:0001706

DOI: 10.1371/journal.pdig.0001706

Access Statistics for this article

More articles in PLOS Digital Health from Public Library of Science
Bibliographic data for series maintained by digitalhealth ().

 
Page updated 2026-09-13
Handle: RePEc:plo:pdig00:0001706