This is a test version of Biostars. For the public version, visit https://www.biostars.org.
GenBank file for BAC clone from Roswell Park collection?

I am working with three BAC clones ordered from BACPAC Chori. I can locate and identify them via NCBI Clone Database but cannot find the Accession Number or GenBank file. I want to upload the file into a program like Genome Compiler to simulate a digest and confirm my clones.

What is the best way to go about this?

bac sequencing

Have you tried to contact CHORI for more information? Can you post an example clone ID?

Hi, thanks for the reply. Sure, one clone ID is RP23-30O22 which can be found here: http://www.ncbi.nlm.nih.gov/clone/609136/ I think I am able to find a GenBank/FASTA file now by clicking on the Clone Placement tab but when I try to import, I get an error saying I must include an ORIGIN tag in my file. Also the sequence seems to only be for the INSERT but there are multiple cut sites in the vector (EcoRI) so there seems to be some guesswork as to which backbone fragment the insert was ligated to. Is there no sequence for the whole BAC clone?

I think the only way to get technical support may be to email the main guy in charge of the whole BAC consortium but I wanted to make sure this wasn't a basic question before resorting to that.

Thank you.

Looks like you don't get the actual sequence but information about the placement of these clones on the assemblies (under the clone placement tab) and then the sequences of the bac ends (under the Associated sequences tab).

If you need the sequence then you could get the sequence by using start-end coordinates in the clone placement tab from appropriate assembly.

Assembly    Asm. unit           Chr  Seq. ID     Start       End         Length  Method   Conc.   Confidence  Placed by
GRCm38.p3   C57BL/6J            13  NC_000079.6 67,319,608  67,516,814  197,207 end-seq     Y     Unique      NCBI
Mm_Celera   Primary Assembly    13  AC_000035.1 69,791,056  69,990,653  199,598 end-seq     Y     Unique      NCBI

Which database would you look at to get that specific section? (Can you specifically download just that segment and not the whole chromosome?) Apologies if this is really simple; I'm very new to learning about all these databases.. Also would you recommend any programs for processing? I'm currently using Genome Compiler but it crashes everytime I try to upload the file which I *think is the 200 kb insert fragment.

You can use Entrez e-utilities (see my answer in this thread: Retrieving batch of selected regions sequences from a NCBI db ).

Here is a URL to retrieve the sequence for GRCm38 for the start-stop (just the BAC segment) in the table above (http://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=nuccore&id=372099097&strand=1&seq_start=67319608&seq_stop=67516814&rettype=fasta&retmode=text ). This will download the sequence as a fasta format file to your local desktop.

id above = gi number which will work for now but NCBI is going to do away with these starting Sept 2016.

0 answers

No answers yet.

Log in to answer this question.