This is a test version of Biostars. For the public version, visit https://www.biostars.org.
What cause the differences between genes annotations from different databases?

I want to get the information of all genes on human Y chromosome, then I found the statistics in different databases --Ensembl (GENCODE), NCBI, HGNC -- are dissimilar.

For example, protein-coding genes numbers:

CCDS 63
HGNC 45
Ensembl 63
NCBI 73

So what leads to these number be different?

By the way, is RefSeq gene data the same as NCBI homo sapiens annotation release?

next-gen gene database

Ultimately HGNC is responsible for all human gene nomenclature. Other databases may add database specific annotation but if you need a list of approved genes for Y chromosome then HGNC is authoritative source.

So other databases will contain all official gene symbols and names from HGNC and add their specific annotations?

Yes but other resources sometimes lag behind HGNC. HGNC occasionally updates official symbols and/or names and the old ones become synonyms and the changes are not always picked immediately by others, it depends on their update cycle.

1 answer

The different resources you cite do different things and are not necessarily in sync. The CCDS tries to identify annotations of protein-coding regions in the human and mouse genomes that are consensual across several groups/institutes. The HGNC is in charge of attributing official names and symbols to genes. NCBI's RefSeq is a collection of sequences that are annotated as belonging to a gene and/or linked to other NCBI resources. Ensembl provides a full genome annotation integrating many information types. Of the resources you cite, Ensembl is the only one that annotates the underlying genome. When doing a bioinformatics project, select one reference and stick to it. Don't mix and match, this would be asking for trouble. I would recommend using Ensembl because it's much better organized and integrated than NCBI resources.

Log in to answer this question.