UniProt has a database of disease-associated and neutral SNPs called humsavar. I used their mapping tool to map UniProt/SwissProt AC identifies to Ensemble transcript and protein IDs. As expected, each UniProt entry has more than 1 match in the Ensemble database. So here comes the question. How do I tell that UniProt variant annotation always references the same protein and not an alternative splice or something? The docs state that all UniProt data are given for a single canonical reference, but I can hardly believe that so many SNPs can be matched without reference ambiguity. I've found that the database itself has 2 different variant categories: natural variants (SNPs) go under VAR IDs and alternative sequences (splices) go under VSP IDs, but there is no references about possible combinations of the two. I'm going to use the database for machine-learning purposes, hence I need to be sure about the reference sequences.
Thank you in advance.
0 answers
No answers yet.
Log in to answer this question.