Univariate-guided sparse regression for Biobank-scale high-dimensional omics data
Joshua Richland,
Tuomo Kiiskinen,
William Wang,
Wenhui Sophia Lu,
Balasubramanian Narasimhan,
Trevor Hastie,
Manuel Rivas and
Robert Tibshirani
PLOS Genetics, 2026, vol. 22, issue 9, 1-18
Abstract:
We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (uniLasso). UniLasso is a two-stage penalized regression procedure that leverages univariate coefficients and magnitudes to stabilize feature selection and produce sparse predictive models. Building on its theoretical and empirical advantages, we adapt uniLasso for application to the UK Biobank, a population-based repository comprising over one million genetic variants measured on hundreds of thousands of individuals from the United Kingdom. We further extend the framework to incorporate external summary statistics via uniLasso ES (external scores). These signals guide the regression toward variants with prior evidence of association by informing penalty weights and sign constraints. Both uniLasso ES and uniLasso ultimately fit multivariate models using individual-level target data; the external statistics guide, rather than replace, this fitting. Our results demonstrate that uniLasso attains predictive performance comparable to standard Lasso while selecting substantially fewer variants, yielding sparser and potentially more interpretable models. Moreover, it remains competitive with other PRS estimation methods, such as PRS-CS and lassosum2.Author summary: The growing ability to measure genetic variation across hundreds of thousands of individuals has created new opportunities to investigate the genetic basis of human traits and disease. One common goal is to construct polygenic risk scores, which combine information across genetic variants to predict an individual’s genetic predisposition to a phenotype or disease. However, widely used statistical methods often identify large numbers of variants, making the resulting models difficult to understand. We implement uniLasso, a sparse regression technique that produces simple genetic prediction models while maintaining strong predictive accuracy. The method uses signals from univariate analyses to guide a multivariate regression model toward a smaller set of variants. We further demonstrate how summary statistics from external genetic studies can be incorporated into the uniLasso framework (through uniLasso with external scores, or uniLasso ES), allowing the model to prioritize variants that have previously shown evidence of association. Applying these methods to large-scale genotype data from the UK Biobank, we find that uniLasso (and its extension with external scores) produces substantially sparser models while achieving predictive performance comparable to widely used approaches. These results suggest that combining simple genetic signals and external information with high-dimensional regression can reduce the complexity and improve the accuracy of polygenic risk prediction.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1012314 (text/html)
https://journals.plos.org/plosgenetics/article/fil ... 12314&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pgen00:1012314
DOI: 10.1371/journal.pgen.1012314
Access Statistics for this article
More articles in PLOS Genetics from Public Library of Science
Bibliographic data for series maintained by plosgenetics ().