Hi Istvan,
Thank you for your comments and kind explanation.
I was going to say the replicates are faithful if the coefficient is high like 0.97.
Now that I read your comments, correlation analysis is not useful for this purpose.
However, I'm still a little concerned about technical variation.
Since the first RNAseq was performed a little different platform (not totally different, as I know, it's different version of Illumina, probably, HiSeq 2500 for the first RNAse, and HiSeq 300 for the second RNAseq), some kind of technical variation affect a lot of counting reads during sequencing process even if it's assumed that there's no biological variation.
This is why I thought I need to do something and prove the replicates is identical.
I happened to find a paper about SERE (simple error ratio estimate). If you know this paper, do you think this paper would be good to show faithful replication?
Thank you, again, for your comments and advice. I learned a lot.. SS
You need to find out the source of the RNA. The question of biological replicates or technical replicates is very important and due to the source material. Where did the (each of six) RNA come from? Your collaborators will know if they are the same cells or different. Pearson correlation is irrelevant for this answer.
Hi karl.stamn
Thank you for the reply and comments.
As you can see some kind of table above. It's biological replicates.
RNAs were prepared from sample1, 2, and 3, and each sample were treated in different condition, like sample 1 was in condition 1, and sample2 in condition 2, and sample 3 in condition 3. Prepared RNAs were used first RNAseq.
And three months later, I did the exactly the same thing. So, in theory, data from sample 1 of the first RNAseq is supposed to be the same as data from sample 1 of the second RNAseq, like this.
Could you let me know why pearson correlation is irrelevant for this?
Is that because of some kind of variation that might cause significant different read counts between replicates?
Then, would Spearman correlation can give better explain between replicates?
Do I have to normalize expected counts to do Spearman correlation? Would just standardization works, instead of normalization?
Thanks, SS
I did Spearman correlation, and I got this results 0.95 (0.95) for sample 1, 0.97 (0.96) for sample2, 0.98 (0.97) for sample3 (EC values (TPM values)) similar to the results of Pearson correlation.
I expected that the Spearman correlation gives higher coefficient because it eliminates variance caused by the differences in read counts.
But the question unanswered is whether this coefficiency 0.95 is enough to convince the replication (Second RNAseq counts = the first RNAseq counts, not exactly same, very similar enough to ignore the minor differences in whole gene expression profiles) or not.
Please, comment and point out things that I miss or misunderstand.
Thanks, SS