This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Multiple Sequence Alignment using muscle

I tried to align 112626 sequences of alfalfa using muscle command, but I get segmentation fault 11 error which is mainly due to low running memory. Then I split into 8 smaller files with 14100 sequences in each file. The muscle command ran smoothly in those files and generated output. However, when I used the -profile option in muscle command to join those 8 aligned files (2 at a time) it gives me error message (* WARNING Invalid byte hex 00 in FASTA sequence data, ignored ERROR * Internal error MSA::ExpandCache, ColCount changed). Why am I getting this error message? and how can I join those 8 small files into single aligned file?

multiple sequence alignment

It seems you want to align a transcriptome or an EST set using muscle, and this makes no sense. Muscle is designed to align sequences with moderate or high similarity (either because they are homologous, or share structural features). What is your goal? Maybe you want first to group the 112626 sequences into clusters with a certain level of similarity, then align these clusters?

While you posted a specific error in this new thread, we have tried to help you with this line of analysis in a past thread: Multiple Sequence Alignment using muscle

What you are trying to do does not make scientific sense.

Hello atitparajuli2018!

Questions similar to yours can already be found at:

We have closed your question to allow us to keep similar content in the same thread.

If you disagree with this please tell us why in a reply below. We'll be happy to talk about it.

Cheers!

0 answers

No answers yet.

Log in to answer this question.