Economics at your fingertips  

Non-parametric estimation of population size changes from the site frequency spectrum

Waltoft Berit Lindum () and Hobolth Asger ()
Additional contact information
Waltoft Berit Lindum: Bioinformatics Research Centre, Aarhus University, C.F. Møllers allé 8, 8000 Aarhus C, Denmark, Phone: +45 87165763
Hobolth Asger: Bioinformatics Research Centre, Aarhus University, 8000 Aarhus C, Denmark

Statistical Applications in Genetics and Molecular Biology, 2018, vol. 17, issue 3, 10

Abstract: Changes in population size is a useful quantity for understanding the evolutionary history of a species. Genetic variation within a species can be summarized by the site frequency spectrum (SFS). For a sample of size n, the SFS is a vector of length n − 1 where entry i is the number of sites where the mutant base appears i times and the ancestral base appears n − i times. We present a new method, CubSFS, for estimating the changes in population size of a panmictic population from an observed SFS. First, we provide a straightforward proof for the expression of the expected site frequency spectrum depending only on the population size. Our derivation is based on an eigenvalue decomposition of the instantaneous coalescent rate matrix. Second, we solve the inverse problem of determining the changes in population size from an observed SFS. Our solution is based on a cubic spline for the population size. The cubic spline is determined by minimizing the weighted average of two terms, namely (i) the goodness of fit to the observed SFS, and (ii) a penalty term based on the smoothness of the changes. The weight is determined by cross-validation. The new method is validated on simulated demographic histories and applied on unfolded and folded SFS from 26 different human populations from the 1000 Genomes Project.

Keywords: Coalescent theory; population size; regularization; site frequency spectrum (search for similar items in EconPapers)
Date: 2018
References: Add references at CitEc
Citations: Track citations by RSS feed

Downloads: (external link) (text/html)
For access to full text, subscription to the journal or payment for the individual article is required.

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link:

Ordering information: This journal article can be ordered from

DOI: 10.1515/sagmb-2017-0061

Access Statistics for this article

Statistical Applications in Genetics and Molecular Biology is currently edited by Michael P. H. Stumpf

More articles in Statistical Applications in Genetics and Molecular Biology from De Gruyter
Bibliographic data for series maintained by Peter Golla ().

Page updated 2021-05-24
Handle: RePEc:bpj:sagmbi:v:17:y:2018:i:3:p:10:n:2