Hi Devon,
Could you elaborate on why you'd expect to see a normal distribution for many genes, given that RNA-seq count data is generally over-dispersed? I am currently analyzing a large RNA-seq dataset with hundreds of individuals, and have not seen an example of a gene with normally distributed RPKMs. Also -- isn't the skew you're describing due to the mean-variance relationship, i.e. greater variance at greater expression values?
Thanks, Allie
I was trying to figure it out myself and just saw this old thread. If still relevant for anybody, different tools assume the distribution either to be normal (limma) or negative binomial (EBSeq and DESeq2). You can find a little bit of explanation here: https://academic.oup.com/bib/advance-article/doi/10.1093/bib/bbx122/4524048