Hi everyone,
I have a couple of questions regarding best practices in enrichment analysis for RNA-seq data, particularly in microbial systems, based on the following paper and post:
Post: Methodological problems are extremely common for enrichment analysis - beware the pitfalls before you publish
Paper: "Urgent need for consistent standards in functional enrichment analysis"
As the paper notes:
"In the case of ORA for differential expression (e.g., RNA-seq), a whole genome background is inappropriate because, in any tissue, most genes are not expressed and therefore have no chance of being classified as DEGs. A good rule of thumb is to use a background gene list consisting of genes detected in the assay at a level where they have a chance of being classified as DEG."
Given that microbial RNA-seq data doesn’t involve tissue-specific expression, I’m curious about the most appropriate approach for defining the background set in this context.
1. Recommended Background for Microbial RNA-seq:
Which of the following would be the best choice for the background gene set?
a. All genes
b. Genes filtered with TPM > a certain threshold (e.g., 10)
c. Genes filtered with CPM > 1
d. DEGs with Log2FC above/below a certain threshold
2. In case of GSEA and ORA, Should DEGs Be Separated by Regulation?
I’ve come across different perspectives on these points, but I’m still unsure about the best approach. Any guidance would be greatly appreciated!
Thanks in advance for your help!
rna-seq