Hopefully after all this manipulation the final files contains reads that are still in sync across both R1/R2 files.
Would be nice if the sequencing centers (or Illumina) would clean this up first, but I doubt that will ever happen.
Sequencing centers generally don't (or more specifically should not) release any data that has issues with it. The continuous downward trend for cutting costs (to give the consumer a better rate) may result in less diligence about the quality of the final data. It is possible that your data came from such a provider. With the billions of reads now possible from each sequencing run, losing a few million here and there should not make a big difference as long as the losses are uniform across entire set of samples.
Quality of samples, library prep method and quality of final libraries can all affect the sequencing, so it is possible that the issue may well be upstream of the actual sequencing.
Scanning and trimming programs (bbduk,cutadapt, fastp etc) have been around for a long while and are able to handle these sorts of issues with the right combination of command parameters. If you are happy with the end result, then that is great but much of this was likely because you were confined to working on Galaxy, where you have much less flexibility about choosing program options.
I suggest that you use this BBMap tool instead https://bbmap.org/tools/polyfilter
Can you comment on what fraction of data has this issue?
I'm going to give that a try, I'm hoping I can use that with a slightly older version of BBduk. With the BBduk I was using the most recent recommended settings that I had seen posted by Brian. I'm thinking it was k-31 with 2 mismatches. This seemed to get rid of a lot of the Poly-G, but with some minor spot checking there were still obvious Poly-g at the ends of the reads. With BBduk I had to assume they were removed. But with Cutadapt using adapterX, it would show if there were any additional preceding C or G present so if shown it I could repeat the trimming.
I'm not sure how to quantify on the fraction, but it's significant.
Even with using BBduk on the adapters I had issues. It was more obvious on my nextera Novaseq 6000 reads with the transverse adapters. Using BBduk with the recommended nextera adapters, with the recommended k-23 settings, it would remove a considerable amount of adapters, then with spot checking there were perfect matches at the ends of some of the reads. I had found that Fastp, cutadapt, trimgalore, BBduk adapters were also removing a signifant amount of biological reads. The Adapterremoval program either using the Illumina Universal or Nextera transverse adapter seemed to work adequate.
What do you mean by this?
That really should not be happening, unless the read remaining after removal of the poly-G becomes so short that it fails the minlength setting. Can you post the complete
bbdukcommand line that you had used for reference. One trick you could use is you can simply provideliteral=GGGGon command line as the sequence to find and remove.I have a feeling that there is something wrong with your data, if you are having so much trouble. Generally the sequence provider should have checked this (especially if a significant amount of reads contain this issue) and not released the data. It sounds like you have data from more than one provider?. If so, is your organism known to have difficult to sequence regions (repeats etc)? This is normal genomic sequence correct?
Sorry, meant Transposase sequence
R1 CTGTCTCTTATACACATCTCCGAGCCCACGAGAC R2 CTGTCTCTTATACACATCTGACGCTGCCGACGA
My test for biological removal from adapter removal. Map reads that map at 100% Match ID, then run adapter removal on those. So this didn't include reads that didn't match at 100% due to minor errors. These are reads that match at 100% for Naegleria Fowleri NF001, Plasmodium Ovale, Pseudomonas A. and Mycobacterium. I didn't take the time to find out which ones were hitting.
My BBduk V.39.08 for adapters was Trim-R, min length 10, k-23, mink-11, hamming distance 1. Using the adapters above. Just a quick peek I could see exact matches to the front of the adapters on both reads from 5bp to maybe 20bp on the ends. I run Adapterremoval with default settings and same adapters removed them, with no adapters showing on fastqc or Falco.
I had just found that I can do the Poly-G Poly-C with Cutadapt using GGGGGGGGGGGGGGGGGGGGX as an adapter with the "X" allowing only removal from the 3' end, or XGGGGGGGGGGGGGGGGGGGG allowing only from the 5' end so it's not trying to find any internal poly-g regions.
This is the cutadapt Poly-g output AFTER running with BBduk poly-G, showing now all the poly-g is gone. This is from Element Aviti.
Here is the cutadapt poly-C first run, showing there are still 12.1% that retains a C base after removal, so I'm running again.
I do realize that I may lose some information by removing all of the poly-G and poly-C from the ends.
I don't think I can use that tool, as I'm using the Usegalaxy.org platform, and unable to use it on my personal computer. I can use BBduk, but it's version is a bit older so I can't directly use the polyfilter.sh tool.
Then consider doing this on a local computer. You are going to have a lot more control over the options/data that way and can properly diagnose issues easily.
I tried setting bbtools up on my computer, but couldn't figure how to set it up to run. I'm currently running a windows 10 computer.
If you’re on Windows, the easiest way to run BBTools is via WSL2 (Ubuntu [a flavor of Linux]). Install WSL following Microsoft’s guide: https://learn.microsoft.com/en-us/windows/wsl/install
This’ll allow you to access a Bash shell inside Linux. Once Ubuntu is installed, you can install BBTools.
Ok, I was completely wrong on this, well partially for sure. I went through my reference sequences that I'm working with and broke them into smaller pieces. All three of my primary sequences Nagleria Fowleri, Plasmodium Ovale, and Tuberculosis BG24 happen to ALL include long stretches of Poly-C and Poly-G. So yes, I was removing a lot of Poly-C/G but was also removing my target.
I'm sure it's likely that there are artefactual Poly-C and Poly-G as well. Since I'm trying to separate sequences that map to my references from human, I think initially I'm not going to worry about these as I will be able to retain these from direct mapping.
Question on the Poly-G Poly-C issues
I'm currently trimming the Poly-G and Poly-C from my reads. I'm using Cutadapt using a 156bp CCCCC....., trimming first from the 5' end, followed by the 3' end. From what it looks like, as the percentage of reads are nearly identical on the R1/R2 reads. I'm assuming that since the percentage of reads with the Poly-C are nearly identical, that they are artifacts??? My assumption is that since Read2 is the reverse complement of Read1, that Read2 ends would be Poly-G if it was biological. I also get similar results by trimming Poly-G tails. Am I reading this correctly??
Poly-C Trimmed from the 5' to 3' end Total read pairs processed: 7,186,087 Read 1 with adapter: 2,604,798 (36.2%) Read 2 with adapter: 2,529,032 (35.2%)
Poly-C Trimmed from the 3' to 5' end Total read pairs processed: 7,185,659 Read 1 with adapter: 2,303,990 (32.1%) Read 2 with adapter: 2,300,002 (32.0%)
Sort of and/or yes. Read 2 is reading the second strand so that sequence is reverse complement with respect to read 1. Unless you have a 300 bp insert and you sequence 300 cycles (exactly the same length as the insert), the two reads (R1/R2) will not be perfect reverse complements of each other. If the reads are pure C's or G's then they would be reverse complements of each other irrespective of the insert length.
Please don't add these "observations" as new answers. This thread has already become long and convoluted. New readers are going to get lost trying to make sense of the flow.
I keep running into issues with these reads, which is causing mapping or assembly issues. So I'm trying different methods to clean adapters and Poly G/C issues. Since I'm giving Cutadapt another try, as the other trimmers were leaving adapters in my trimmed reads. So I'm separating the reads that had adapters 1. at the 3' end, 2. at the 5' end, and 3. where no adapters were found.
I had forgotten to mention, these specific reads (Element Aviti) were ones where no adapters were found nor trimmed. If I trim Poly-G or Poly-C tails from either the 3' or 5' ends, and if nearly the same percentage was trimmed from each read, my thought is they are artifact. Do you agree???
On my reads where adapters were trimmed from either the 3' or 5' ends, the end where an adapter was trimmed should be biologic, and the end where no adapter was found MAY have an artefactual Poly-G or Poly-C??
The adapter issues I found was that TrimGalore was leaving adapters. Mostly I read that I should trim adapters at the 3' end, which TrimGalore probably did fine. But after a manual search for adapters and their reverse complement, I find there are reverse complement adapters at the 5' end of reads. From what I have read to trim these reverse complement adapters is to use the reverse complement of Read1 adapter at the 5' end of read 2. But what I was finding were the reverse complement of read1 adapter at the 5' end of read 2, but I was also finding the reverse complement of Read 2 at the 5' end of Read 1. I had not seen this documented anywhere. I also found that neither FASTQC or FALCO was not identifying any of the reverse complemented adapters, so the adapters looked clean, which they were not.
Due to the way libraries are made, in Illumina sequencing there should be no adapter sequence on 5'-end of reads. Way sequencing primer works, it always reads into the first base of the insert for both ends. So any adapter sequence should only be present on 3' end of reads. (see illustration in Figure 1 here: https://knowledge.illumina.com/library-preparation/general/library-preparation-general-reference_material-list/000003874 )
Sadly since we can't access your data, our ability to help you has reached a dead end. As I said before, you really should consult with an expert, who can directly access your data and help address the questions you are encountering. It is one thing to describe (e.g. in this thread) things you are "observing" and quite another to have actual access to the entire dataset first-hand.
It was my understanding that with shotgun sequencing that sequencing doesn't necessarily start at either end so it can start from 3' to 5', which would end up as a reverse complement adapter in the reads. I'm only using full length adapter sequences on my trimming to prevent issues of using the first 13bp.
As confirmation, since I know at least several of the strains that match at 100% ID, after a trimming operation I map the reads to only allow 1bp error (not including N's). Since I get additional reads that Map, it's obviously taking away adapters or artifacts.
Either the watson or crick strand can be on top in shotgun libraries but sequence is always reported 5' --> 3' direction.
If crick strand is sequenced then the sequence is reported in the 5' --> 3' direction which is reverse complement of watson strand.
I suspect that if you have any shotgun reads to test with, if you setup Cutadapt using the reverse complement of your adapters as a 5' adapter with standard settings, you will find the same adapters I'm talking about.