Thanks so much for this, Phil. I didn't think that adapter sequence can be anywhere in the read. I thought it's always at the end of the read, since the GBS technique we used 1) cuts up the DNA at specific places using a cutter enzyme and then 2) attaches the adapters to those cut sites. Is there something I'm missing?
My reads were (fortunately) very high quality so there's not much chance the reads would be majority Ns. Can you please give another example for what could've happened?
If the Sequence length distribution data is post trimming (cutadapt is pre-trimming) and if you expect ~40 bp to be removed then the peaks seems to have shifted correctly in trimmed data?
Thanks for your reply! I was under the impression that the Trimmed Sequence Length plot is showing the number of nucleotides that are trimmed per read, which would imply that it's trimming, in some cases, >100 bp, which doesn't make sense if I'm only asking it to trim ~40 bp total, per read. What do you mean by "cutadapt is pre-trimming"?
I wonder, does the x axis label "Length trimmed (bp)" mean the base pair at which the read was trimmed? Which I think is similar to your idea?
I don't use cutadapt so I am not so familiar with its metrics. I thought that the plot you were showing us in figure 1 was data prior to trimming. Is that not the case?
This must show the plot of reads (length) that remains after trimming. Not the bases that got trimmed.
Frankly, I'm not certain what the "Trimmed Sequence Length" plot is showing. I can't figure out another explanation besides that, as you suggested, this plot shows read lengths before trimming, and the other plot shows lengths after trimming. However, there's a huge peak around 140 in the second plot that I can't explain...