EconPapers    
Economics at your fingertips  
 

An interpretable LightGBM model for predicting coronary heart disease: Enhancing clinical decision-making with machine learning

Lang Deng, Kongjie Lu and Huanhuan Hu

PLOS ONE, 2025, vol. 20, issue 9, 1-26

Abstract: Background: Coronary Heart Disease (CHD) is one of the major burdens of cardiovascular diseases worldwide. Traditional diagnostic methods, such as coronary angiography and electrocardiogram, face challenges including high costs, subjectivity, and high misdiagnosis rates. To address these issues, this study proposes a prediction framework for CHD based on the LightGBM algorithm, aiming to improve the accuracy and interpretability of CHD risk prediction. Methods: This study utilized three publicly available datasets: BRFSS_2015, Framingham, and Z-Alizadeh Sani. The BRFSS_2015 dataset was used for model training, while the Framingham and Z-Alizadeh Sani datasets were employed for validation. Data preprocessing included cleaning, feature engineering, and handling missing values. The LightGBM model was selected for its efficiency and performance, and SHAP (SHapley Additive exPlanations) values were used to enhance model interpretability. Model performance was evaluated using metrics such as accuracy, precision, recall, F1-score, and AUROC. A CHD scoring system was developed based on the model’s predictions to assist clinicians in risk assessment. Results: The LightGBM model demonstrated excellent performance, achieving an accuracy of 90.60% and an AUROC of 81.06% on the BRFSS_2015 dataset. After parameter tuning, the model’s accuracy improved to 90.61%, and the AUROC increased to 81.11%. On the Framingham dataset, the accuracy improved from 83.96% to 85.26%, and the AUROC increased from 62.86% to 67.37%. On the Z-Alizadeh Sani dataset, the accuracy improved from 78.69% to 80.33%, and the precision increased from 74.40% to 76.36%. Conclusions: SHAP analysis revealed that age, smoking status, diabetes, hypertension, and high cholesterol were the most influential features in predicting CHD risk. The developed CHD scoring system provided a user-friendly tool for clinicians to assess patient risk levels effectively.

Date: 2025
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0330377 (text/html)
https://journals.plos.org/plosone/article/file?id= ... 30377&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pone00:0330377

DOI: 10.1371/journal.pone.0330377

Access Statistics for this article

More articles in PLOS ONE from Public Library of Science
Bibliographic data for series maintained by plosone ().

 
Page updated 2025-09-20
Handle: RePEc:plo:pone00:0330377