Thanks John! I have already compared my two in vivo phenotypes and the results are indeed quite nice, but we wanted to compare to in vitro conditions to see what stress factors might be associated to each of the two in vivo conditions, hence the published data come handy.
You are absolutely correct on the differences between our technology, data quality and all the statistical noise - it's almost like comparing apple to pear, but getting some ballpark figure will be nice enough for us.
One question regarding TPM. I have been doing some readings (including the Wagner paper) but I don't quite understand what Z (mean read length) stands for? Is it referring to the average bp of the reads from the sequencer? Or is it referring to the average bp mapped to each gene in a given sample? Will greatly appreciate it if someone can clarify for me. Thanks again!
RPKMs are terrible for statistics. Would it be possible to just analyse your samples and compare the resulting fold-changes/DE genes to those from the published study? That might yield nice results.
I thought about doing FC as well but have another dilemma. What would you use as a cut-off? Presumably I will have to use some arbitrary cut-offs, e.g. if >2 log2FC then highly abundant. Is there any way to make it more objective?
Also, to do FC, will you:
Thanks again!
I am facing a similar situation with data on GEO only being present in log2 FPKM. After you converted to TPM, how did you end up testing for differential expression?
Thanks!