PS, here are stats for Evigene versus Transdecoder, computing proteins for a known gene set, the Arabidopsis plant. You can see why I recommend above using Evigene's computed ORFs, which are a better match to these carefully curated plant proteins. You can reproduce this test fairly easily, or with another reference species, and may get other results if you use other options/data.
Published gene set: Araport11_genes.201606.cdna and Araport11_genes.201606.aa, n=48,359 transcripts/proteins of 27,655 gene loci Arabidopsis thaliana Genome Annotation Official Release, date: June 2016
Method of ORF from cDNA transcripts:
- Evigene $evigene/scripts/cdna_bestorf.pl -nostop -cdna Araport11_genes.201606.cdna -outaa arath16ap_evgorf.aa
- TransDecoder.LongOrfs -t Araport11_genes.201606.cdna -m 30
- TransDecoder.Predict -t Araport11_genes.201606.cdna
Predicted ORFs versus published Araport11_genes.201606.aa
- Evigene cdna_bestorf
Total=50,846, Identical=46,639, Missed=23, aveSizeDiff= +1.3, sumSizeDiff=67573
(includes 2,510 "UTRorf" extra orf/transcript) - TransDecoder.LongOrfs Total=492,753, Identical=41,233 Missed=25, aveSizeDiff= +3.7, sumSizeDiff=181235 (includes ~10 ORF per transcript)
- TransDecoder.Predict Total=77,414, Identical=28,399, Missed=1363, aveSizeDiff= -0.4, sumSizeDiff=-19799 (includes 29,055 "UTRorf" extra orf/transcript)
For those Missed, there are some very short "proteins" in this published plant gene set, including one with 1 amino only !, the others with a few aminos up to 20 or so .. these are below the computed ORF cut-off of 30 aa.