This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Search Motif In Raw Reads

Hi,

How can I find a specific pattern (~20nt long) and its reverse complement from my raw reads (paired-end data so 2 fastq files) ? and then extract them into 2 new fastq files ?

Thanks a lot,

N.

motif search read

And what does this pattern look like? How specific - like no mismatches/indels?

4 answers

Use Biopieces www.biopieces.org) and try something like this:

read_fastq -i file1.fq,file2.fq | patscan_seq -i -c -p acgtactagctagctactagc[2,1,1] | grab -p PATTERN -K | write_fastq -o matched.fq -x

This allows for 2 mismatches, 1 insertion and 1 deletion (and ambiguity codes).

gunzip and paste both fasts on the fly, with awk , group by '4 lines', test if the pair matches your needs (here first read starts with the regex "^ATGGG[GC]AAAA*" split the output into two files using awk.

paste <(gunzip -c input_1.fastq.gz) <(gunzip -c input_2.fastq.gz)  |\
awk '{a[i]=$0;++i;if(i==4){split(a[1],S,"\t"); if(S[1] ~ "^ATGGG[GC]AAAA*") {for(j=0;j<4;++j) printf("%s\n",a[j]);}; i=0;}}' |\
awk -F '    ' '{print $1 >> "select_1.fq"; print $2 >> "select_2.fq";}'

You can try meme:

http://meme.nbcr.net/meme/cgi-bin/meme.cgi

Although if you are using raw reads, you'll probably have quite a large signal-to-noise ratio. It might be a better idea to align them first and then use the regions that they align to as inputs to a motif finding program.

It's in the idea to search for a specific fusion bewteen two genes so I prefer first to extract only the reads with a particular motif (the fusion point from gene A) and then align against the genome.

Well, MEME might not be the right choice for this, but FIMO could be used... It allows motif discovery based on PWM .

Log in to answer this question.