This is a test version of Biostars. For the public version, visit https://www.biostars.org.
RNA-Seq, why estimate the distribution of your data?

Hello

I am new in the field of bioinformatics and have a background in biochemistry.

When doing differential expression analysis on RNA-seq data, edgeR and DEseq2 estimate the distribution of your data. Why? They both do it, so i guess it is important, but i have no clue why they do it.

Thanks in advance.

rna-seq edger deseq2

1 answer

Typically there are relatively few replicates, so it's difficult to accurately estimate the variance (or dispersion if you prefer that term) on a per-gene level. Consequently, both tools pool information across genes to get more robust measures. Since the reliability of a p-value is determined in large part by how accurate you've estimated variance, this ends up being a major benefit.

Of course if you have a lot of samples (e.g., a thousand) then this isn't really needed.

Log in to answer this question.