Thanks for your input! This analysis is in the discovery phase. After seeing difference in cell population, I ran DE within those separate populations and still got 1-3 DEGs. I still put them into GSEA and got some enriched pathways and currently trying to check out the leading edge genes. Sometimes I bump into a gene for which the expression is elevated by only one sample. I also plot genes that fell into the enriched pathway on a heatmap and sometimes see within group heterogeneity (e.g. in my disease group 1-2 samples express these genes highly, while others only modestly), so I think there is within group heterogeneity.
Hi guys! Im quite new to scRNAseq.
I'm working with human patients data: 4 controls and 11 disease samples. I annotated the broad cell types and then subclustered cell types of interest and annotated their subpopulations.
Then I did sample level pseudobulk and PCA. The samples didnt separate by group and I noticed they were separting by sex genes. After removing sex genes and repeating PCA the samples still didnt separate by groups. I then correlated the PCs with sample metadata and found that several PCs correlated with abundance of some subpopulations.
I proceeded with DESeq2 indicating group and sex in the design. I practically got no significant DE genes (occasionally 1 or 2, but nothing particularly interpretable). I did DE analysis on all cells and then within subpopulations only if enough cells were available.
I also plotted pseudobulk PCA using 200 and 500 of identified DE genes but still didnt see clear group separation (PCA attached).
I fed the DE genes to GSEA and got some significantly enriched pathways. For some of them there is quite plausible biological explanation stemming from histological analysis. I also looked at the leading edge genes and plotted them for some extra reassurance.
For one cell type I also noticed that one sample contributed to these pathways due to extreme phenotype and removed this sample, which left that pathway at FDR 0.053
My main question is: Does this workflow seem appropriate so far?
Also how would you normally take pathway level results further? Im not sure how to move beyound this pathway is enriched and seems interesting. In general, what usually follows?
Thanks in advance for your input
1 answer
There's a clear separation in this PCA. But that's not surprising given you filtered for that specifically. For bulk DE, if you noticed there's difference in cell population %, can you run DE separately for the cell types? E.g. if your population has neutrophils and macrophages, calculate DE for disease vs control Macrophages
To take it farther, it would depend on your specific research questions, and what stage of the project this is coming. If this is support or discovery for example. In discovery, you may consider experimental manipulations, if possible, or go back to other data and see if it was missed previously. In support, you just want to show explicitly the connection with previous data.
Also depends on what pathways are enriched. I would start by examining the expression level of key genes in the pathway. Is it a subset of cells that really express it, or are they scattered throughout? What are the most co-variable genes? Can these be associated to a feature within the disease population?
Log in to answer this question.