This is a test version of Biostars. For the public version, visit https://www.biostars.org.
How to extract equal number of R1 and R2 read from raw Illumina data

Dear All,

Somehow, We got different number of R1 and R2 reads in Illumina paired end sequencing data. Is there any way to extract equal number of R1 and R2 reads from these raw files. these are just pair end library not strand specific.

[root@psgl data_new]# grep -c '^@'  SS_5W_R1.fastq

26623063

[root@psgl data_new]# grep -c '^@' SS_5W_R2.fastq

25803102

[root@psgl data_new]# grep -c '^@' SS_7W_R1.fastq

42474961

[root@psgl data_new]# grep -c '^@' SS_7W_R2.fastq

41089376

Thanks

rna-seq

If you want to use this method always include a few characters that follow @ sign (which are generally the machine serial) in line 1 (e.g. grep -c "^@M1023" file_name).

2 answers

Quality score strings can contain or start with "@" so this is not a reliable method. Please use "wc" instead, or use an actual bioinformatics tool to count the reads and test formatting.

Thanks a lot.

with this command, it is coming correct number:

awk '{s++}END{print s/4}' fastq file name

Thanks again!

cat SS_7W_R1.fastq | echo ((`wc -l `/4))

Log in to answer this question.