EconPapers    
Economics at your fingertips  
 

A spectral framework for measuring diversity in multiple sequence alignments

Vaitea Opuu

PLOS Computational Biology, 2026, vol. 22, issue 9, 1-17

Abstract: Machine learning (ML) methods for proteins and RNAs rely on multiple sequence alignments (MSAs) and related datasets such as experimental mutagenesis libraries, yet the amount of usable information they contain remains unclear. Here, a spectral measure of information is recast into an interpretable quantity for MSAs, denoted Leff, defined as the number of fully independent alignment positions that reproduce the observed sequence diversity. Applied to RNA MSAs, this measure shows that evolutionary constraints nearly halve diversity relative to the secondary structure alone, quantifying functional and phylogenetic restrictions beyond base pairing. The same analysis indicates even lower effective diversity in proteins, reflecting tighter packing and coevolutionary constraints. Leff further correlates with protein structure prediction accuracy, anticipating cases with insufficient evolutionary signal. When applied to experimentally and computationally generated libraries, it measures both produced diversity and cross-library overlap, quantifying novelty rather than redundant sampling. Together, these results establish Leff as an operational tool to estimate effective information in MSAs, anticipate modeling difficulties, and guide protein and RNA design.Author summary: Machine learning has transformed biology, predicting protein structures, uncovering evolutionary rules, and designing new RNA and protein sequences. Almost every such method learns from large collections of related sequences, and the field largely assumes that more data means better models. But more is not always richer. A collection of thousands of sequences may hold far fewer independent evolutionary signals, because so many entries are near-copies shaped by shared ancestry or by designs that scarcely depart from a single template. We rarely know how much real information a dataset carries, let alone how to measure it. Here I introduce the effective length, a simple measure of the genuinely independent information in a sequence collection. It reveals that natural RNA and protein families are far more constrained than their size suggests, anticipates how reliably their structures can be predicted, and distinguishes new sequence libraries that add real information from those that merely repeat what we already know.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014778 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14778&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014778

DOI: 10.1371/journal.pcbi.1014778

Access Statistics for this article

More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().

 
Page updated 2026-09-27
Handle: RePEc:plo:pcbi00:1014778