Hi, I am using limma for differential gene analysis of RNA seq results. I have encountered the following issues: I cannot find any genes with significant padj values (less than 0.05). Although some genes have extremely high logFC values. (I use TPM values for difference analysis)
But when I use DESeq to perform differential analysis on the same sample (using counts), I can normally obtain a set of differentially expressed genes.
I want to know what is causing this situation. All help will be greatly appreciated!
I will attach the limma code I used below.
1 answer
There is a lot non-standard here. First, if you say TPM then this is RNA-seq most likely, so normalizeBetweenArrays() as normalization is odd since it's meant for arrays, not counts. Then TPM by itself is not a good choice for DE analysis (many threads on it, please google to learn more). Also, limma for RNA-seq if using pre-normalized counts typically uses the trend-pipeline, see limma vignette. Since you seem to have counts available why not just following either the DESeq2 manual, or the limma (limma-voom) manual to do the analysis rather than this custom approach you do here (which in part is wrong I would say).
This is indeed RNA seq data, as I mentioned earlier. Regarding why standardization is carried out, I believe that TPM is aimed at standardizing the effective length of genes, while the parameter "normalizeBetweenArrays()" is aimed at standardizing between different samples (eliminating batch effects). I think there is a difference between the two. Secondly, I carefully reviewed the results of my limma differential analysis and found that using a p-value (<0.05) for filtering can obtain differential genes, with approximately two thousand different genes. But we got nothing by filtering through Padj. Now I believe that some reason is causing my Padj value to be abnormal. But I am not good at statistics.
Log in to answer this question.