Hello everyone,
I am currently studying a research paper in which the author integrated datasets from multiple public databases and their own samples. I located the data using the SRA numbers provided by the author, but I found that the raw data amounts to over 800 GB. The supplementary materials from the author state that all data were processed uniformly in the upstream analysis. As a self-learner, I don't have the resources to handle such a large volume of data; I only have access to my local computer.
Luckly, the author provided GEO accession numbers in the main text, and each GEO dataset includes expression profile data. However, a new challenge has emerged. The author combined multiple datasets, some of which provide raw counts values, while others provide TPM values, and some even include counts values that has already been normalized. Based on my current understanding, if I want to conduct differential analysis, I need raw counts values. So, I am wondering if it is feasible to download only the non-raw data datasets from SRA and process them into count data myself. I am concerned that inconsistent data processing methods might lead to discrepancies in the results. But this seems to be the best approach I can think of.
I would greatly appreciate it if any experts could advise me on whether this method is reasonable. Alternatively, if anyone has better suggestions or alternative approaches, please share them with me. I am eager to learn and improve.
Thank you in advance for your help :)
data-integration
differential-analysis
transcriptomics
data-normalization