Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
Gul Rukh Khattak,
Konstantinos Patlatzoglou,
Joseph Barker,
Libor Pastika,
Boroumand Zeidaabadi,
Aidan R Birdi,
Jiayu Huo,
Ahmed El-Medany,
Hesham Aggour,
Yixiu Liang,
Antonio H Ribeiro,
Jeffrey Annis,
Antonio Luiz Pinho Ribeiro,
Junbo Ge,
Daniel B Kramer,
Jonathan W Waks,
Evan Brittain,
Nicholas Peters,
Fu Siong Ng and
Arunashis Sau
PLOS Digital Health, 2026, vol. 5, issue 9, 1-21
Abstract:
Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Augmented Patient Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,269), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining, discourages learning of spurious cohort-specific features, and instead promotes clinically meaningful variability within cohorts. This leads to improved out-of-distribution robustness, with gains of 9–40% in downstream label prediction performance. This work provides insights into pretraining strategies for more clinically deployable and generalisable foundation models.Author summary: Artificial intelligence (AI) can learn from large collections of electrocardiograms (ECGs) and support the development of new diagnostic tools without relying on extensive manual annotation. However, it remains unclear how the data used to pretrain these models influences their ability to generalise across different patient populations.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001623 (text/html)
https://journals.plos.org/digitalhealth/article/fi ... 01623&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pdig00:0001623
DOI: 10.1371/journal.pdig.0001623
Access Statistics for this article
More articles in PLOS Digital Health from Public Library of Science
Bibliographic data for series maintained by digitalhealth ().