Thank you so much Michael! You were right, it just went through if I did adapter trimming after UMI extraction.
Sometimes, a break from the usual routine is what that works for a dataset!
Thanks again, Rituriya.
Hi All,
Background: I have completed adapter trimming and checked QC on Illumina NextSeq miRNA single end reads of length 75bp. I want to run umi_tools to extract the UMI information before I align the reads to the reference. I am unable to run umi_tools extract.
Command used:
umi_tools extract --stdin=XYZ_R1-trim.fastq.gz --bc-pattern=NNNNNNNNNNNN -L XYZ-extract.log --stdout=XYZ-UMIextracted.fastq.gz
I have a 12 bp UMI barcode here. I think the pattern could be the culprit here. I am new to this UMI analysis, could anyone please share their insight as to what is the mistake here? Error message seen on screen:
Traceback (most recent call last):
File "/home/xyz/.local/bin/umi_tools", line 11, in <module>
sys.exit(main())
File "/home/xyz/.local/lib/python2.7/site-packages/umi_tools/umi_tools.py", line 57, in main
module.main(sys.argv)
File "/home/xyz/.local/lib/python2.7/site-packages/umi_tools/extract.py", line 330, in main
new_read = ReadExtractor(read)
File "/home/xyz/.local/lib/python2.7/site-packages/umi_tools/umi_methods.py", line 971, in __call__
umi_values = self.getBarcodes(read1, read2)
File "/home/xyz/.local/lib/python2.7/site-packages/umi_tools/umi_methods.py", line 726, in _getBarcodesString
umi_quals = [bc_qual1[x] for x in self.umi_bases]
IndexError: string index out of range
Also, am I supposed to use whitelist command before extract? This is not single cell RNA data and hence I omitted that step.
Thank you.
Hi Rituriya,
You named your input "trim"; are all reads long enough to extract the 12 bp sequence? Try to run the UMI extract before trimming.
Cheers,
Michael
Thank you so much Michael! You were right, it just went through if I did adapter trimming after UMI extraction.
Sometimes, a break from the usual routine is what that works for a dataset!
Thanks again, Rituriya.
In Python, a string is a single-dimensional array of characters. The string index out of range means that the index you are trying to access does not exist. In a string, that means you're trying to get a character from the string at a given point. If that given point does not exist , then you will be trying to get a character that is not inside of the string. Indexes in Python programming start at 0. This means that the maximum index for any string will always be length-1 . There are several ways to account for this. Knowing the length of your string (using len() function)could certainly help you to avoid going over the index.
Log in to answer this question.
Hi again,
If I followed UMI extraction first and then adapter trimming later, there is a weird long trail of N's (approx 20 N's) towards all the read ends. So to see if something has changed drastically due to UMI extraction, I did adapter trimming in a more effective manner and then did UMI extraction.
Cutadapt command:
UMI extraction command as mentioned in original post but using these adapter-trimmed reads as input.
Now following is my situation in question:
Now I do not see too many trailing N's in my data but this PhiX contamination level is bothering me. Please help me understand this behaviour. Due to UMI extraction I know the reads have become shorter but this is too drastic and tells me something is not right with UMI extraction.
Note: There is a suspect of higher PhiX contamination during the library prep and hence first step to check levels.
Thank you, Rituriya.
Please start a new thread if you have a new question.