Thank you very much for giving such in-depth answers Istvan Albert and Mensur Dlakic! I have actually also looked for some numbers on how many phage sequences are even in the db and was wondering about those small numbers. I am also using the new "nt_viruses" databank. As I have just recently started working in this field I thought this would be reasonable to choose. I do not know what your opinion is on this db?
This is the description of the db: "The Viruses nucleotide collection consists of GenBank+EMBL+DDBJ+PDB+RefSeq sequences, but excludes EST, STS, GSS, WGS, TSA, patent sequences as well as phase 0, 1, and 2 HTGS sequences and sequences longer than 100Mb. The database is non-redundant. Identical sequences have been merged into one entry, while preserving the accession, GI, title and taxonomy information for each entry."
Also I have decided to not filter out environmental & uncultured samples as I realized that this also filters out some interesting hits, that, after doing some research in the GeneBank, were not even uncultured and environmental samples. So I would have lost some info here. But actually I also had some hits on the Caudovirales sp. as you mentioned @mensur. As you mentioned this does not give much information but I will have to see what I make from that. They are not the only hits with high qcov and %identity for my sequences.
But you made me realize that these cutoffs, I read about so often, are not to just be copied but rather I should estimate which parameters make sense for my individual case!
The bacteria you mention are common contaminants in labs and reagents...
Just getting a hit on a phage itself is not informative enough. The query coverages and identities matter a lot. Case point I am investigating a e-coli phage contamination and the phage that I found is only about 50% similar to known phages, and I get lots of other hits to different species as well. Once assembled half the phage is practically unknown.
Is there some sort of generally used/accepted cut-off? In the literature many papers use values like at least 50% identity of 90% query coverage.
Those numbers 50% and 90% feel very "human" oriented rather than fact-based.
Seem like numbers that are easy to remember and strong enough if you can get them. The problem is that your phage might not match anything at those criteria.
The novel phage genomes I am finding are the other way around, more 50% query coverage and 90% identity over those regions :-)
Long story short, I think the phage diversity is far larger than anticipated and assigning a phage to a species based on a partial match is less reliable than one would expect. Out of curiosity, I looked at the RefSeq representative viral database
shows
filtering for phages:
produces
Thus about a third of all viruses blast knows about are some sort of phage.
But then there are almost a million prokaryotes in RefSeq alone, so knowing about just 6 thousand phages seems like a substantial underestimation.