Thanks Carlo for the suggestion.
I have one more question after reading this.
How the low expressed gene (or zero counts) is related with type 1 error and false discovery rate. Is there a mathematical correlation?
Please suggest.
Thanks
Hi all,
I have a question regarding the filtering of lowly expressed genes in analysis of RNAseq data. What I understood from literature is that these low counts genes are basically a noise and not a true picture of differentially expressed genes, so they need to be removed if very low counts are observed for all the samples (not only in one sample).
I wonder if there is any other basis of it, mainly in terms of :
1). Differential expression
2). Statistical
3). any Molecular / Biological
I appreciate any suggestion.
Thanks
Ankit
The main reason behind the idea of discarding low count genes is to NOT test genes for which we presume a difference in expression would not be relevant. If you do less tests, then the correction for multiple testing becomes less stringent, and the overal power of your experiment increases. This concept, called 'independent filtering' is well explained in this publication: Independent filtering increases detection power for high-throughput experiments
Thanks Carlo for the suggestion.
I have one more question after reading this.
How the low expressed gene (or zero counts) is related with type 1 error and false discovery rate. Is there a mathematical correlation?
Please suggest.
Thanks
Log in to answer this question.
StatQuest gives some nice details on this.