Hello Dave, Thanks for your valuable comment, I have run the script to retain the longest isoform and I have obtained the following, what do you think about this?
I appreciate your comment again.
################################
## Counts of transcripts, etc.
################################
Total trinity 'genes': 277062
Total trinity transcripts: 277062
Percent GC: 41.59
########################################
Stats based on ALL transcript contigs:
########################################
Contig N10: 3086
Contig N20: 2125
Contig N30: 1538
Contig N40: 1097
Contig N50: 794
Median contig length: 370
Average contig: 603.55
Total assembled bases: 167220426
#####################################################
## Stats based on ONLY LONGEST ISOFORM per 'GENE':
#####################################################
Contig N10: 3086
Contig N20: 2125
Contig N30: 1538
Contig N40: 1097
Contig N50: 794
Median contig length: 370
Average contig: 603.55
Total assembled bases: 167220426
Since you used
trinitythis must be RNAseq data. In that case getting many contigs is not unexpected nor is some "redundancy". Did you run BUSCO in transcript mode?Hello Geno,
If it is RNA-seq data and I ran BUSCO in Galaxy in transcriptome mode:
A version of the genome already exists, however, the authors have not yet authorized its use for massive studies: