The above experiment the replicates should have been termed technical replicates. There are always going to be biological differences due to the nature of biology (more like noise), but the linked paper was showing the impact of technical variability on the outcome of the experiment. The yeast were genetically homogenous, so really the only major source of variance would have been plating effects or other batch effects. This is the whole point of a cell line, if you can't assume that two cultures of a cell line as supposed to be biologically equivalent, then why have cell lines. The only reason there will be differences is due to variance in how the protocol was performed. The only biological difference was the single deletion of a gene, aside that differences withing genotypes are either natural noise or variance introduced by the experimentalist.
Say I had an experiment where I compared infected and uninfected cells from a cell line. If I had 3 replicates of each condition per timepoint, these would be more like technical replicates, NOT biological replicates. The cells and agent are the same, so the reason for having replicates is to determine how the protocol may have added variance to the experiment. In this case I need to use replicates to determine/mitigate any impact variances in performing the experiment (e.g. had to use the restroom during the experiment) will have. In other words, the replicates are there to capture/mitigate how the experimenter may have induced variance into things.
If I'm getting PBMCs from infected and healthy patients, the situation changes. If I have three sick patients and three healthy ones, I have three biological replicates. These replicates allow me to capture the impact a given individual will make. In other words, I can tell if a response might be general to all infected individuals, or if it might be due to something common to only one of the three people.
However in this case, because I still have to collect blood, process it and so on, there's plenty of risk for problems that might lead to bias through variance in performing the protocol. So although I'm able to capture my biological variance, I am still missing the technical variance. In this case the best thing would be to perform the processing/extraction/etc three times on a single blood draw.
I don't think these differences are always explained clearly.
I think I'm going to tape that preprint (and your article) to the door to our core facility.
Yeah, we hope this article gets more attention. There needs to be more studies like these for other applications.
Towards the end, the article mentions that advances like paired-end and longer reads can improve performance and alleviate some of the problems by reducing the incidence of these.
It seems like the idea in the end was to reduce the 'risk' of having problematic samples by increasing the number of replicates, however, just like better technology doesn't totally negate these risks, neither does increasing the number of replicates.
I guess the best approach then is to use better sequencing technology with increased numbers of replicates. If one doesn't have the money to do this, which ends up being the better route: better sequencing tech or more replicates? Which ends up being more cost effective?
Can it be assumed that paired-end and longer reads will have the same dynamics when dealing with "bad" replicates?
I noticed that in the w/t samples that the "bad" replicates seemed to have occurred within a more or less contiguous region. I wonder if the authors had balanced their culture plate(s). Hopefully they didn't just use a 96 well plate with one side having w/t and the other having their deletion.
I am one of the authors on the paper. It's great to see this discussion, shame I'm three months late.
Sequencing technology doesn't obviate the need for biological replicates; variability is a fundamental feature of biology and needs to be measured in all experiments. See this paper. The comment regarding paired-end data is that we may have been able to remove some of the artifacts if it wasn't SE data; no guarantee that it would 'rescue' the bad samples, however.
No we didn't use a 96-well plate. All 96 cultures were grown separately and the libraries were prepared in batches of 24 with 12 of each condition per batch - randomly assigned. The bad replicates were not consistent with batch or lane.