This is a test version of Biostars. For the public version, visit https://www.biostars.org.
DNA Motif discovery from a large number of FASTA sequences

Hello

I wish to find the recurring/enriched motif(s) in a large number (approx. 7000) of DNA FASTA sequences. I have given MEME a try but it does not scale as intended.

Any thoughts/suggestions/recommendations for another software or strategy are highly appreciated.

Thanks. FG

sequence motif alignment

Hi, sorry for bothering you but do you want novel DNA motifs for example from ChIP-Seq studies?

Hi. Yes. I am looking at discovering novel motifs from a RNA+Chip-Seq study.

Hello, did you attempt to use MEME via command line or the web version of it? In the second case, I would just recommend switching to command line.

I see, then I think you should be just fine with the MEME suite, what is your problem exactly?

2 answers

Try Homer, http://homer.ucsd.edu/homer/motif/

Thanks EagleEye for the suggestion. However, with Homer, what would you recommend for a 'background' FASTA file.

2. Background Selection findMotifs.pl/findMotifsGenome.pl

If the background sequences were not explicitly defined, HOMER will automatically select them for you. If you are using genomic positions, sequences will be randomly selected from the genome, matched for GC% content (to make GC normalization easier in the next step). If you are using promoter based analysis, all promoters (except those chosen for analysis) will be used as background. Custom backgrounds can be specified with "-bg <file>".

Hey if you are looking for motif enrichment use centrimo or MEME-ChIP, because there you can upload you sequences from DEG and unchanged genes and look for Transcription factor binding sites, for example or for miR datasets. If you are using MEME or any othe above mentioned tools, you can use the out put for further categorical analysis, as to which genes possess these motifs and their recorded possible function by using GoMO

Log in to answer this question.