This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Estimating Effective Population Size (Ne) From Rnaseq Data

Hello, we're getting RNAseq data from increasing numbers of "emerging" model species for which little is known. The data often wasn't meant to provide population genetic estimates, but could perhaps be used to provide some.

Is it possible to estimate Ne from the SNP patterns that can be identified from:

  • a single. sequence pooled from many individuals from a single population?
  • a single diploid genome?
  • RNAseq data where multiple individual from a single population (siblings or not) were independently sequenced.

Cheers, yannick

A somewhat related question, but specific to pools is here.

population next-gen sequencing

Could you please elaborate on the reasoning? Why do you think it is possible, in principle, to estimate the effective population size from RNAseq data?

I am not an expert, but I think the biggest issue you will have is that to have absolute Ne estimates, you need SNP data from the neutrally evolving part of the genome. If one does whole genome sequencing and then SNP calling, most of the SNPs are going to be neutral or nearly-neutral. If you are sequencing RNAseq, a lot of the SNPs you gather will be under purifying selection, and the Ne estimates, which are based mostly on allele frequencies, will have an unavoidable ascertainment bias. Someone else from the related question you pointed might be able to give more details.

1 answer

It seems that people use RNA-seq or exome capture data for demographic inference by using synonymous mutations, thereby getting around the issue of non-synonymous mutations being non-neutral. Therefore, it should be similarly possible to use your RNA-seq data to estimate Ne (your third point above) if you use the synonymous mutations only?

For example, see:

Fraïsse C, Roux C, Gagnaire P, Romiguier J, Faivre N, Welch JJ, Bierne N. (2018) The divergence history of European blue mussel species reconstructed from Approximate Bayesian Computation: the effects of sequencing techniques and sampling strategies. PeerJ 6:e5198 https://doi.org/10.7717/peerj.5198

Log in to answer this question.