EconPapers    
Economics at your fingertips  
 

Biomarker discovery and patient stratification in pancreatic cancer using incomplete multi-omics data

Alejandra Paja-García, Rafael Romero-Becerra, Tero Aittokallio and Alberto López

PLOS Computational Biology, 2026, vol. 22, issue 9, 1-29

Abstract: Pancreatic ductal adenocarcinoma (PDAC), with a 12% 5-year survival rate, is the most aggressive type of cancer. Early diagnosis for this pathology is rare, and conventional treatments such as surgery, radio- or chemotherapy, have little to no effect on reducing mortality. Machine learning (ML) approaches could be used to identify biomarkers that help clinicians stratify patients and improve treatment outcomes. However, most ML techniques perform poorly with incomplete data, which is usually the case in real-world settings, often forcing researchers to discard valuable information. In this study, unsupervised ML algorithms capable of dealing with missing modalities were applied to incomplete multi-omics data from PDAC patients to identify clinically meaningful patient subgroups. Through a large-scale clustering benchmark including six omics layers, we discovered two novel subgroups with statistically significant differences in survival and recurrence after surgery, particularly within the first two years, when most patient deaths occur, as well as distinct tumor mutational burden. Comprehensive multi-omics analyses revealed substantial molecular differences between patients in both groups, identified three methylation biomarkers to stratify patients, and highlighted dysregulation in key oncogenic pathways. Importantly, the identified groups are different from previous PDAC classifications, both in their patient composition, prognosis, and in the oncogenic gene pathway profiles exhibited. Using an independent cohort, we further demonstrated that both the prognostic value of these subtypes and their underlying biological characteristics are reproducible. These results could lead to better stratified treatment regimens to improve the prognosis of PDAC patients.Author summary: Pancreatic cancer has one of the highest mortality rates among all cancer types, and it is difficult to find common molecular characteristics across tumors that can be used to create new treatments for patients. Machine learning algorithms are been used to analyze the molecular data of cancer patients and identify similar groups, for example, patients that share the same mutations in specific genes. This information can then be used to create tailored therapies for individuals and obtain the best patient outcomes. We applied machine learning algorithms to molecular data from pancreatic cancer patients (e.g.,: DNA, mutations, proteins) and identified two novel groups, with one of the groups showing lower survival probability, and an increased risk of cancer recurrence compared with the other. Further analysis revealed key cellular pathways and that three molecules could be used to separate patients into the two groups. This opens avenues for therapies personalized for every individual, to ensure that pancreatic cancer patients receive the best care they can.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014735 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14735&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014735

DOI: 10.1371/journal.pcbi.1014735

Access Statistics for this article

More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().

 
Page updated 2026-09-13
Handle: RePEc:plo:pcbi00:1014735