Hello Kevin,
Thank you for your answer.
I agree with you, for DEA in RNA-Seq, the preprocessing consisting in the suppression of lowly expressed genes only seems now a gold standard and adapted to the normalization step that follows.
My concern is rather why a person would perform preprocessing on RNA-Seq data for DEA as we mentionned and preprocessing on methylation data for DMR with more extensive criteria. I imagine that there are people processing the 2 types of data in different ways. As you mentioned, this is probably because of the nature of the data: the ranges of beta values and normalized counts are completely different, as their distributions, so it can make sense to use different approaches.
Good point, I will check what kind of distribution except the methods for tumor deconvolution, I do not know yet, I just started on this question. Could you recommend other types of transformations (in addition to Z-scores) to try?
Finally, could you please explain why do you think Z-scores are particularly interesting in the field of deconvolution?
Thank you, Jane
Sorry, I just saw your answer. Yes I mean identification of cell populations in bulk by tumor deconvolution.
Using pure cell populations is an interesting approach, used in supervised methods, as CIBERSORT, EPIC, xCell, ... on gene expression data. I assume that when using supervised approaches, the "cleaning of the dataset" might have less importance than when using unsupervised approaches. I can use both approaches in parallel, but my focus here is mainly on unsupervised methods. That is why I would like to start with a clean and meaningful matrix.
Z-scores are intuitive. What I cannot figure out is (if/why) their use might improve in some way the analysis, besides the interpretation.