EconPapers    
Economics at your fingertips  
 

A Bayesian framework for multivariate differential analysis

Marie Chion and Arthur Leroy

PLOS Computational Biology, 2026, vol. 22, issue 9, 1-27

Abstract: Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.Author summary: In many areas of biology, researchers need to decide whether one condition truly changes the levels of proteins or other measured features. Standard statistical tools often reduce this question to a single significance value, which can be hard to interpret and may handle missing measurements poorly. In this work, we developed a probability-based method for comparing groups that focuses on two practical questions: how large is the difference, and how certain are we about it? Using quantitative proteomics as our main example, we show that this approach can deal naturally with uncertainty, make use of relationships between related measurements, and, in simple cases, avoid unnecessary filling-in of missing values. Because the method relies on formulas that can be computed directly, it remains fast enough for large studies while giving results that are more informative than traditional testing alone. Rather than encouraging yes-or-no decisions based on a threshold, our framework helps researchers judge the size and reliability of observed changes. We hope this offers a clearer and more useful way to study biological differences in proteomics and beyond.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014637 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14637&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014637

DOI: 10.1371/journal.pcbi.1014637

Access Statistics for this article

More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().

 
Page updated 2026-09-06
Handle: RePEc:plo:pcbi00:1014637