Hi people!
I am analyzing Ribo-seq libraries prepared with the QIAseq miRNA Library Kit and sequenced as 75 bp single-end.
The expected RPF length is approximately 28–34 nt. However, most reads show this structure:
[short GC-rich sequence][QIAseq 3′ adapter][12 variable nt][AGATCGGA...]
Adapter sequences:
3′ adapter: AACTGTAGGCACCATCAAT
5′ adapter: GTTCAGAGTTCTACAGTCCGACGATC
The 3′ adapter is present in approximately 96% of reads, but the 5′ adapter is absent in both orientations.
Among 2,495,800 reads, the sequence length before the 3′ adapter is:
<17 nt: 1,845,987
17–24 nt: 393,075
25–27 nt: 62,838
28–34 nt: 64,816
>34 nt: 33,272
Only 2.6% of reads have a 28–34 nt insert. These inserts are also highly GC-rich, with a mean GC content of approximately 78%.
My questions are:
Is the QIAseq 5′ adapter normally absent from Read 1? Are the 12 variable bases after the 3′ adapter the UMI from the RT primer? Could the short GC-rich inserts be rRNA/tRNA fragments or indicate a problem with size selection or library preparation?
I would appreciate a QIAseq Read 1 structure diagram or advice from anyone who has processed this type of library outside the QIAGEN pipeline.
1 answer
Taking them in order: yes, yes, and probably rRNA.
The 5' adapter not showing up is expected. On this kit Read 1 starts at the 5' end of the insert, so the 5' adapter sits outside what actually gets sequenced. Your observed structure (insert, 3' adapter, 12 nt, AGATCGGA...) is exactly what it should look like, and yes, those 12 bases are the UMI.
The bigger thing is that you're doing Ribo-seq on a miRNA kit. Its size selection is tuned for ~22 nt miRNAs, not 28-34 nt footprints, so you're preferentially keeping fragments much shorter than the ones you want -- that's what the 74% under 17 nt is. And 78% GC in short fragments is a very rRNA-shaped signature. Ribo-seq is dominated by rRNA without a dedicated depletion step, and a miRNA kit doesn't include one.
Easy way to confirm: bowtie the short bin against an rRNA/tRNA reference and see what fraction maps. I'd expect most of it.
Your 64k reads in the 28-34 window are probably real footprints, but 2.6% is a rough yield. Deduplicate on the UMI first and see how many unique molecules are left -- that's the number you actually have to work with, and it'll be lower than 64k.
Log in to answer this question.