This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Biopython Translate With N In The Sequence

Hi. I have the following sequence:

CAGGTGCAGCTGGTGCAGAGCGGCAGCGAGCTGAAGAAACCTGGCGCCTCCGTGAAGGTGTCCTGCAAGGCCAGCGGCTACACCTTCACCAGCTACGCCATGAACTGGGTCCGCCAGGCCCCAGGCCAGGGACTGGAATGGATGGGCTGGATCAACACCAACACCGGCAACCCCACCTACGCCCAGGGCTTCACCGGCAGATTCGTGTTCAGCTTCGACACCAGCGTGTCCACCGCCTACCTGCAGATCTGTAGCCTGAAGGCCGAGGACACCGCCGTGTATTNNTGTGCGA

There are a couple of N's in there. I would like to use biopython's translate function on the seuqence, but this throws the following error: "Codon TNN is invalid"

Is there a way to get this function to return a default amino acid such as 'X' when the translation is unsuccessful? Any ideas?

biopython

1 answer

I don't have any problem to translate your sequence using biopyhton

>>> from Bio.Seq import Seq
>>> dna = Seq("CAGGTGCAGCTGGTGCAGAGCGGCAGCGAGCTGAAGAAACCTGGCGCCTCCGTGAAGGTGTCCTGCAAGGCCAGCGGCTACACCTTCACCAGCTACGCCATGAACTGGGTCCGCCAGGCCCCAGGCCAGGGACTGGAATGGATGGGCTGGATCAACACCAACACCGGCAACCCCACCTACGCCCAGGGCTTCACCGGCAGATTCGTGTTCAGCTTCGACACCAGCGTGTCCACCGCCTACCTGCAGATCTGTAGCCTGAAGGCCGAGGACACCGCCGTGTATTNNTGTGCGA")    
>>> dna.translate()
Seq('QVQLVQSGSELKKPGASVKVSCKASGYTFTSYAMNWVRQAPGQGLEWMGWINTN...XCA', ExtendedIUPACProtein())

The TNN codon is valid, and it's translated to X, just as you suggested.

Cheers!!

AHHH, Looks like my problem was my alphabet. I hadn't noticed that I was using UnambiguousDNA. Using generic_dna fixed the issue

please mark the answer as accepted if you think that it's correct. You and the person who answered will receive extra score.

Log in to answer this question.