I'm not sure about the overall average gene in the assembly because I'm only particularly interested in 87 genes. So for the 87 genes the average length would be around 600bp. And yes, I did encountered some of the genes that were spread across 2 contigs, causing only partial sequence of the genes can be obtained.
Can you elaborate further on how do I proceed with blastn or blastx? Because as I've mentioned, I've tried doing reference-guided assembly instead and have extracted the unmapped reads to find out what were they. But most of the reads were unclassified. I did blastn and blastx on these unclassified reads but no significant similarity were found.
BUSCO result for the de novo assembly was actually quite good:
C:90.1%[S:87.4%,D:2.7%],F:5.6%,M:4.3%,n:446
402 Complete BUSCOs (C)
390 Complete and single-copy BUSCOs (S)
12 Complete and duplicated BUSCOs (D)
25 Fragmented BUSCOs (F)
19 Missing BUSCOs (M)
446 Total BUSCO groups searched
It would be much easier to help you if you provided more details. A strain of what? What reads did you have and how were they assembled? What is the depth of coverage?
What is the assembly size when you throw out contigs smaller than 2K, 3K, 5K? You may already have a proper genome size once you throw out the small stuff.
An Eimeria tenella strain. It was illumina paired-end DNA reads with depth of coverage at 80x. With de novo assembly, MEGAHIT and SSPACE were used, both with default parameters. While reference-guided assembly, BWA MEM was used.
I'm not sure about throwing out the contigs smaller than 2k, 3k and 5k, because the average contig size was only 1402.