Yes, when I do that it is missing sequences. If I wc the original input and the parallel output, it goes from 10000 to 9588. In the parallel output, the first line starts as:
>CTTCACTAGCT
>r9
sequence
whereas in the original file, it started as:
>r1
sequence
>r2
sequence
etc.
I think there are other instances of missing sequences (where each block is made) but it is hard to find them without going through 10000 lines manually. Do you have any ideas how I could trouble shoot this? Parallel is extremely useful to me but with this little issue I cannot use it.
edit: I have found another instance where there is a skip in the read numbers and where the header is altered.
>r187.1 |SOURCES={GI=330827700,fw,273802-273903}|ERRORS={27:C,30:C,32:T,62:A,96:A,99:A}|SOURCE_1="Aeromonas veronii B565 chromosome" (db14d9defaae9617cf9b20eb8bb2b46eefae8000)
CGGGATCGTGCGGGTGGCTCTGTGCATCCTCGTTGGTTTGAGCGGGGGATGAGTCTGCCGTCAGTGCAGTGGGCCAGAGCAACACCCCGCCAAGCAACAAG
>CE_1="Aeromonas veronii B565 chromosome" (db14d9defaae9617cf9b20eb8bb2b46eefae8000)
ACAGGTCGTGGTTGTGCAGCCCGCCAGCAGCAATTCGAGAGTCATGGGACGCCCCACTATGATGGACGCTCCCACCACGCCAGCATGCAGACCGTGCATCT
>r195.2 |SOURCES={GI=330827700,bw,2584285-2584386}|ERRORS={36:G,40:A,42:A}|SOURCE_1="Aeromonas veronii B565 chromosome" (db14d9defaae9617cf9b20eb8bb2b46eefae8000)
ACCCGCAGCCGACCAAATTGCAGCTGCTGGGGAACGGTGCAGATGGCTACCGGTTGGGTTGATGCTTGCGGCTGCTGGCGAAACTGCACCCGAAGTCGGGC