This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Disambiguation of Non-Unique Clinvar IDs

We are using nucleotide changes to assign unique ClinVar IDs for variant curation and identification of clinical significance. However, when we search for duplicates from ClinVar (via hgvs4variation2.txt) we found that ~2% of ClinVar ID's are non-unique and map to multiple Nucleotide Changes. For example, ClinVar ID 1333944 corresponds to at least 5 different nucleotide changes.

Is this shared ClinVar ID by design since I thought ClinVar IDs were meant to map 1:1 to a particular variant?

I have tried searching for duplicates looking at information in other columns (other than nucleotide change) and still have about the same number of duplicates, is there another variable to include with ClinVar ID to disambiguate these overlapping entries?

Is there another database that might be able to allow better map 1:1, ID:Variant Nucleotide Change along with at least as much information on clinical significance as ClinVar?

vcf snp clinvar

0 answers

No answers yet.

Log in to answer this question.