No unfortunately the sequence as I have shown is not in one line and I want to convert that to a one line sequence only contains A, T, C, G and N
Hi
I have a fasta file started by
>chr1
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
I want a fasta which is a one line character string; just keep the nucleotides characters like 
Basically I should remove anything that is not T, C, G, A or N. After replacing any such characters with "N"
I have tried this but gives an empty file
cat input_fasta.fa | sed -r 's/[RYKMSWBVHD]/N/g' > output_fasta.fa
Can you help me?
Thank you so much
2 answers
wget http://hgdownload.cse.ucsc.edu/goldenPath/currentGenomes/Homo_sapiens/chromosomes/chr2.fa.gz
zgrep -v ">chr" chr2.fa.gz | tr -d '\n' | sed -e '$a\' > chr2.fa
Do you want to remove the header line in your input file? If so, your output file is not in fasta format. If your sequences are already in one line, the below command can do the trick.
grep -v ">" input_fasta.fa | sed 's/[RYKMSWBVHD]/N/g' > output_fasta.fa
Do you know how to run a PERL script? If so, you can use my script to convert your fasta file to a one line format. Then use the above grep & sed command to do other stuffs. You may download the script @ https://github.com/Qiongyi/custom_PERL_scripts/
For your task, the following command should work:
perl linker.pl input_fasta.fa input_oneline.fa
grep -v ">" input_oneline.fa | sed 's/[RYKMSWBVHD]/N/g' > output_fasta.fa
Thank you
The link says not found
GO TO https://github.com/Qiongyi/custom_PERL_scripts/ AND THEN FIND linker.pl
Log in to answer this question.
Linearize your fasta file using @Pierre's code (which you can easily find by searching for "linearize fasta", should be first hit). Then remove the first column to leave just the sequence.
input:
output: