Hi all, thanks for the comments and answer.
jomo018, since it's barcoding, I don't think I'm using any of the standard alignment utilities you're referring to--could you give some examples?
Also, I'll add the following details which were asked for, in case someone finds them useful:
- Nextera Indices Kit was used, with i5 and i7, for multiplexing.
- The primers we use also have their own old barcodes, apparently.
- By "tags"--maybe these aren't tags per se, but there is sometimes TCAT occurring before the forward primer, and GGAG occuring before the reverse primer.
Here is what I got from BBDuk:
Input is being processed as paired
Input: 178704 reads 51661997 bases. KTrimmed: 23191 reads (12.98%) 605133 bases (1.17%) Total Removed: 2 reads (0.00%) 605133 bases (1.17%) Result: 178702 reads (100.00%) 51056864 bases (98.83%)
Here is from BBMerge:
Pairs: 89351 Joined: 23783 26.617% Ambiguous: 28013 31.352% No Solution: 37555 42.031% Too Short: 0 0.000%
Avg Insert: 357.7 Standard Deviation: 12.2 Mode: 365
Insert range: 52 - 425 90th percentile: 365 75th percentile: 365 50th percentile: 365 25th percentile: 339 10th percentile: 339
Could you elaborate more on the library was prepared?
Have you tried to scan the data with a trimming program? I suggest
bbduk.shfrom BBMap suite. You may have inserts that at smaller than the length of sequencing. While you are at it you could also usebbmerge.shfrom the same suite to see what you get in terms of merging of R1/R2 reads.It sounds like your amplicon library was constructed by the standard Illumina method (i.e., adaptor ligation) and sequenced with standard Illumina (adaptor) primers. If so, then you'd expect a 50/50 mix of amplicon orientations. But @WouterDeCoster is correct, we'll need more details about library prep (e.g., what are the short tags to which you refer) to help you parse the data.
The primer sequences you use in the sequencing step, use the adaptors you link to your fragmented DNA or cDNA)
And the joining of these adapters to these pieces of DNA is fully random (don't get into consideration direction) excepting when you are using a stranded transcriptomic protocol
If using genomic sequences, I am not aware of a protocol that will allow you to get directional libraries, though
Just plain (multiplex) PCR based enrichment & library prep can be directional.
So what are the other methods of amplicon library prep? what if I do not want the 50/50 mix of amplicon orientation?
PCR-based methods (as opposed to ligation) will produce directional libraries. You can either incorporate the Illumina adapter sequences into your amplicon primers, or add them via two rounds of PCR (first round with amplicon primers, second round with Illumina adapters + amplicon overhang).