Thanks seta! I have access to an HPC at my University. For that above analysis I ran with 20 cores with 19GB of ram each. It took a week total. Also, I forgot to mention that the analysis was ran on a unique file that I merged from several samples processed in the same run. Total of 23 libraries (4 fastq files each line paired end; so total of 8 files actually) and an overall total of 184 fastq files (~1 Gb each).
A) Could the fragmentation in your experience be due by the merging of the sample?
B) Should I run a distinct assembly for every single library?
C) Or should I just I considered one total file merged and run the normalization with -min_kmer_cov 1 (default) like you said?
Note: The libraries are from the same tissue but total different treatment if that matters at all.
Thanks!
That's a lot of transcripts. Try using the read normalization parameter. What species are you assembling?
Thanks for your answer. I know it is a lot and I did not run the digital normalization, maybe I should have considering that I have close to a billion raw reads. Do you think that might be the reason of such low N50? Anyway the species is an Hawaiian Squid (no genome annotation at all since it has tons of STR), predicted genome of about 3.8 GB.