How to check Fasta file ASCII characters and fix encoding errors?
I tried building a diamond database but got this error.
Error: Error reading input stream at line 180825: Invalid character (ASCII 0) in sequence
How can I fix it? Is there a tool that checks for this and either repairs or removes the fasta record?
• 2,888 views
•
link
2 answers
Plain text:
sed -i 's/[\d0]//g' xxx.fasta
Gzipped:
gzip -cd xxx.fasta.gz | sed -i 's/[\d0]//g' | gzip -c > clean.fasta.gz
• 0 views
•
link
I read and save in biopython to remove all the trash. Something like:
less biopy_reformat_fastq_remove_short.py
## Colin, Feb. 2018
## Remove short length sequences reported in the fastq (intended for Pacbio downsampling)
from Bio import SeqIO
import sys
good_seqs=[]
c=0
if len(sys.argv) <= 2:
print "Enter input and output files \neg. python biopy_reformat_fasta.py input.fa output.fa"
else:
for record in SeqIO.parse(open(sys.argv[1], "rU"), "fastq"):
# Change the minimum read length
minimum_length = 3000
if len(record.seq) >= minimum_length :
record = record.upper()
good_seqs.append(record)
else:
c = c + 1
print "Found short sequences:" + str(c)
output = open(sys.argv[2], "w")
SeqIO.write(good_seqs, output, "fastq")
output.close()
• 0 views
•
link
Log in to answer this question.
but anyway, you should be very suspicious about the file. It's probably corrupted