EconPapers    
Economics at your fingertips  
 

Large language models enable prognostic stratification of cancer patients using real-world clinical notes

Niklas Kiermeyer, Tim Lenfers, Amin Dada, Julian Friedrich, Sameh Khattab, Eric Knop, Jan Egger, Markus Pauly, Andreas Jung, Grégoire Montavon, Jens T Siveke, Marcel Wiesweg, Stefan Kasper, Ulf P Neumann, Frederick Klauschen, Sylvia Hartmann, Martin Schuler, Philipp Keyl, Jens Kleesiek and Julius Keyl

PLOS Digital Health, 2026, vol. 5, issue 7, 1-18

Abstract: In medical documentation, vast amounts of unstructured text are generated that are still underutilized in current prognostic models. We investigate the potential of self-hosted large language models (LLM) to extract clinically meaningful, patient-specific information from routine clinical notes for personalized risk stratification in cancer care. We collected real-world medical notes from 2,708 non-small cell lung cancer (NSCLC) patients and 814 colon cancer patients documented before treatment at a large comprehensive cancer center. LLMs extracted key prognostic indicators, including comorbidities, metastatic sites, and qualitative descriptors of patient condition, in a zero-shot manner without prior task-specific training. Integrating these LLM-derived features into machine learning models significantly improved the prediction of overall survival compared to TNM staging alone (C-Index: NSCLC, 0.72 vs 0.64; colon cancer, 0.70 vs 0.59), and surpassed models using text embeddings. Based on the LLM-informed risk scores, patients were stratified into four distinct risk groups, enabling reclassification of 61.4% of NSCLC and 68.3% of colon cancer patients. Analysis of model drivers revealed that LLM-derived factors, such as the physical condition, substantially modulated the prognostic impact of TNM stage. These findings highlight the potential of self-hosted LLM to derive prognostically relevant information from unstructured clinical documentation and support clinical decision-making.Author summary: In routine clinical care, medical staff document large amounts of patient information, but only a fraction is captured in structured electronic health records, while much remains in free-text clinical notes. However, most established scoring systems do not use this information and instead rely on a small set of structured variables, such as tumor stage. In this study, we investigated the use of large language models (LLMs) for the extraction of prognostic information from clinical notes without task-specific training. Applying LLMs to the medical records of more than 3,500 patients with lung and colon cancer, we extracted patient characteristics, including mobility impairment, pain, dyspnea, and comorbidities. These features were validated against structured EHR data and expert physician annotations. The LLM-extracted features substantially improved machine learning–based survival prediction and patient stratification beyond conventional measures such as tumor stage. Furthermore, allowing the LLM to derive its own summary score of patient condition provided strong predictors of patient outcome. These findings demonstrate that artificial intelligence can unlock prognostic information from clinical records at scale, supporting more informed and personalized clinical decision-making.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001546 (text/html)
https://journals.plos.org/digitalhealth/article/fi ... 01546&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pdig00:0001546

DOI: 10.1371/journal.pdig.0001546

Access Statistics for this article

More articles in PLOS Digital Health from Public Library of Science
Bibliographic data for series maintained by digitalhealth ().

 
Page updated 2026-07-12
Handle: RePEc:plo:pdig00:0001546