This is a test version of Biostars. For the public version, visit https://www.biostars.org.
don't know exactly how to find my reference genome

Hi I am very new to bioinformatics analysis. I am trying to map a whole genome in galaxy, but I don't know exactly how to find my reference genome (chlamydomonas reinhardtii). I do know It is available on NCBI but don't know which file should be used as a reference genome. I mean: is the reference genome only one single file, or is it multiple files for each chromosome? I'd be grateful if you could help me.

genome assembly sequencing

Thanks Juanjo. I will do that. Much appreciated!

2 answers

You can find the representative genome page for your organism at NCBI (LINK). Actual genome sequence is in this file which is in fasta format. This is the file you should use as a reference.

If you need specific help with Galaxy then consider posting questions to their help forum: https://help.galaxyproject.org/

First, thanks very much for your help. Much appreciated. So, the reference genome is a single file.

Now, I have another question. This .fna file contains ~500,000 letters, while chlamydomonas reinhardti's genome is 110 Mb. How can this reference cover all the genome?

Yes, you're right. Thanks very much for your kind help.

In general, I would recommend you making a BLAST search in the ncbi online BLAST tool. Take one of your contigs and select around 10.000 bps and make a search. Sometimes there is no standard reference genome but someone already sequenced your species. And sometimes what you are assembling is not exactly what you think it is. The BLAST results include a link to the complete sequence that you can download.

Log in to answer this question.