This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Fastq Format Issue

I am doing some experiment with BowTie. Now, I want to do experiment with 150 bps read length. So, I download it from here. And converted to fastq format. Now, I see, the fastq format looks like,

@ERR103405.1 M10_151:1:2:12250:1321 length=302 ATTTACTGCCTTGTGTCTCCAGTGCGCTGAAAATACCTTTATCTTGAAATAAGTTAACTAACTCTTGGATACCTTTAATTAATGCTGGGTTACCACCAGAAATTGTAACGTGGTTAAATAAATCGCCACCAATACGTTTTAATTCATCATAGAACAGCTGGATGTGATTATCGCTGTAGCTGGTGTGATTCTGCATTTACTTGGGATGGTAGTGCTAAAGGCGATATAAAACTCATGACCGCTGAAGAAATTTATGATGAATTAAAACGTATTGGTGGCGATTTATTTAACCACGTTACAAT
+ERR103405.1 M10_151:1:2:12250:1321 length=302 CCCFFFFFHHHHHHHIHJJJJJIIJJIJJIJJJIIGJJJJIIGIJJHIGIIJJIIIJIIJJIJEIJIJFIIIFJGHHGHHFFFFFFFEDCCACCDA?ABDDDDDDCDC@?<ABBBDDDDEDDDC<?B?@BDDDDDB>CC@C:>AADDCACDB@CFFFDDHHBFHEHIIIIIGJIHHEGHIIHE1C?D?GGGIIIIGIFI>BHHIJ@3CHBDGGICHGEHIIGHE>BEDEDE;ACCDDCCA?B=BBCDCCCC@@>>C@CDC>@DCDCDDD<<@?AC(2??BDBDBCDCDDCC::?881<?C>:

Now in NCBI, they described it as "DNA for paried end (150bp) sequencing on an illumina MiSeq". But here it looks it is 302 bps read. Can anybody help me why it is given in above sequence, "length=302" while it is written in the page that it is a 150 bps read.

bowtie genome fastq

2 answers

If you are converting from SRA to FASTQ using the SRA Toolkit, you need to split joined reads. This option puts forward + reverse reads into one file:

fastq-dump --split-spot myfile.sra

And this one generates 2 separate files for forward + reverse:

fastq-dump --split-files myfile.sra

See my blog post for some more details.

Thanks a lot. This is what I really want.

Hi. The link you provided, it clearly says that reads are joined. So you can parse the file and separate the reads till 151 bp each. If you will click on "Metadata" (at link you mentioned), its written that actual_read_length:151.

Yah. I just saw it. Thanks. But, I want 150 bps single read .sra file. Finding no where.

For the same or any file having single end reads?

Any file, better if Human/ChIP DNA.

But should have 150 bps single end reads.

read1 and read2 are in paired end

Log in to answer this question.