Hi,
I have Genebank Protein accession numbers and trying to identify the genes that code for these proteins. I am working with Toxoplasma gondii RH and here are some of my accession numbers: TGRH88_000010 TGRH88_000020 TGRH88_000030 TGRH88_000040 TGRH88_000050 TGRH88_000060 TGRH88_000070 TGRH88_000080 I only know how to code in R and BioMart does not seem to contain my organism. Anyone know how I can accomplish this?
1 answer
These are not genbank accession numbers. They appear to be locus tags submitted with the whole genome shotgun assembly of Toxoplasma by the submitters.
If you look at Toxodb page: https://toxodb.org/toxo/app/jbrowse?loc=CM023083%3A148147..151114&data=%2Fa%2Fservice%2Fjbrowse%2Ftracks%2FtgonRH88&tracks=gene&highlight=
the gene name seems to tbe the same as the protein codes you have above.
Not sure what you are looking to do but your options are going to be limited.
You may need to use a different assembly or map your identifiers to assembly available at NCBI: https://www.ncbi.nlm.nih.gov/datasets/gene/GCF_000006565.2/
Log in to answer this question.