genome sequence
Hi friends,
I found 7 genome sequences for one strain of E. coli in IMG/M. these sequences are different in Genome Size and Gene Count. what can be reason for these differences as they are all for one strain?
bacterial
genome
sequence
• 1,914 views
•
link
updated
by
Mensur Dlakic
•
written
by
Rob •
0 answers
No answers yet.
Log in to answer this question.
More posts like this
-
The genome survey map has miscellaneous peaks
written by zhangtianyuan •In the genome Survey, There are some miscellaneous peaks in the low depth sequences. Comparing these sequences with nt database, they are not found in …
-
galaxy prinseq trimmming sequence
written by Rob •Hi friends, I run prinseq in galaxy with same parameters for my file with 6 sequences twice: why results is different? this is the main …
-
genome compare to the exons
written by Rob •Hi friends how you answer this questions? genome compare to the exome? my answer: We should count the reads against the exons of the gene. …
-
Are housekeeping genes in E. coli essential.
written by Researcher •Hi, I thought that the housekeeping genes that are used for MLST typing of E. coli are essential. But, I found that it is not …
-
Strain specific primers from multifasta file
written by rob •Hi, I have a collection on isolates we are going to inoculate in plants. To quantify them in planta, we want to design strain-specific primers …
-
Add string to list of protein sequences in fasta file with different lengh
written by Jason •How can I do this using python or shell command: Suppose I have 200 protein sequences in fasta file named 'proto.fasta' with different lengths I …
-
Generate local blast database with RefSeq bacteria AND taxonomy
written by even.s.riiser •Dear all, I would like to be able to create my own custom local blast database, as this may be relevant in many different situations …
-
How to know if a genome is curated?
written by Carlos CaicedoDear all I am going to work with E. coli k-12 sub-strain mg1655. In the PATRIC website I found more than 10 genomes available. How …
-
What are the types of variants Next Generation Sequencing technologies can find?
written by davehresearch •By using NGS technology, genetic variants between the patient's DNA sequence and the reference genome can be found. Are these variants same as mutations or …
-
Are there any Bacterial Drug Resistance Databases
written by logust79 •<p>Hello,</p> <p>I'm here to ask whether there's a good bacterial strain drug resistance database available? I know the existence of 'The Comprehensive Antibiotic Resistance Database'. …
literally anything. Library preparation, sequencing technology and the assembly strategy. Do you have a link for each of the six?
thank you Andres What link do you mean? I have this screenshot of my result from IMG/M
https://ibb.co/Tb0YK3m
Escherichia coli str. K-12 substr. MG1655star (E. coli) (University of Oklahoma): Sequencing technology 454, Assembler Newbler, cov 75X link
Escherichia coli K-12 subMG1655 (Pacific Biosciences): Sequencing technology PacBio, Assembler Celera, cov 99X link
Escherichia coli K-12 subMG1655 (The Genome Institute at Washington University): Sequencing technology Illumina, Assembler Velvet, cov 70X link
Escherichia coli K12- MG1655 (Broad Institute): Sequencing technology Sanger-Illumina, Assembler manual processing(?????), cov NA link
Escherichia coli K12- MG1655 (Broad Institute): Sequencing technology PacBio-Illumina, Assembler AllPath, cov 92 link
Escherichia coli K12- MG1655 (Weill Cornell Medical College): Sequencing technology PacBio RS, Assembler NA, cov NA
I would say that the differences in the genome size are caused by the different combination of assemblers and sequencign technologies. Same for the gene count, different annotation algorithms give you different results.
Thanks for great answer. what biological reasons can be the driver of differences?
Which strain would that be? They may have different level of completeness. Some of them may include plasmids as well.
Thanks Mensur for responding. This is E.coli K12 MG1655 .
A couple of these assemblies are incomplete, so it is easy to understand why they are different. For the rest, it would have to be something in annotation procedures, because they shouldn't be that different. I would download them and annotate in an uniform way, and I suspect they wouldn't be so different. If they still are, it is a valid question whether these are in reality the same strain. One presumes that they are, but strain differences would be the easiest biological explanation for why the assemblies are different.