Problem solved, thanks, had never seen WES data before and it was mis-annotated as RNAseq when it was WES.
I'm currently analysing some public RNAseq data, and I've come across this weird coverage profile for one of the datasets.
The reads produce this weird bell curve-like peak over exons, but the coverage extends over the exon boundaries as you can see. There are almost no splices in a file containing almost 70M PE 150 reads. Example given for PHGDH.
Does anyone know what might cause this?
Protocol for library prep as follows...
"Total RNA was extracted and purified from fresh frozen tissues using the TRIzol® reagent. mRNA was purified from total RNA using poly-T oligo-attached magnetic beads. Fragmentation was carried out using divalent cations under elevated temperature in First Strand Synthesis Reaction Buffer (5X)."
Despite polyA selection, there appears to be some 5' bias in the samples.
Thanks
1 answer
Kinda looks like exome capture or DNA contamination… definitely rather strange almost suspicious.
Log in to answer this question.
5' bias is the least of your problems, the coverage itself is super puzzling.
The peaks look too nice and smooth.
It almost seems like the data was mistakenly fitted with a peak predictor of some sorts that turned sharp exon boundaries into an exponential decay shape.
If I took a normal RNA-seq then gaussian kernel smoothed the coverages, I could turn the abrupt exonic coverages into what you show above