Hello,
I'm trying to perform enrichment analysis on 40 proteins that were measured from a 92 protein assay to better understand these 40 proteins.
I understand the bias that can be introduced by using the whole genome background. However, I am wondering whether using the 92 proteins as the background is appropriate from a methodological and statistical perspective. Not surprisingly, none of the tools I have tried (clusterProfiler, MetaScape, gProfiler, ShinyGO) identified any enriched terms when using the 92 proteins as the background.
I would appreciate any ideas, input, or settings to try out in this case.
Thank you!
1 answer
In my head and hands, enrichments are only useful when you have large numbers of genes from (largely) untargeted assays, and you want to generate initial working hypothesis for downstream interpretation. A 92 protein assay must have had a hypothesis when it was designed. Probably some "housekeepers" are in there for normalization, and then the rest to answer a specific question. So I don't really see the point of enriching a predefined set of few proteins. Just intersect the significant ones with the pathway database and see which processes have some intersections. Ignore the stats, they're useless in overrepresentation analysis anyway (in my opinion, due to redundancy, underpowered test, excessive multiple testing burden etc). Or give it to you favorit AI and let it suggest a story based on canonical genes. From there, use your knowledge of the system and see what the literature so far knows about these proteins.
Log in to answer this question.