This is a test version of Biostars. For the public version, visit https://www.biostars.org.
variable length barcodes in SPRITE protocol

the SPRITE protocol (Split-Pool Recognition of Interactions by Tag Extension) was published in Quinodoz et al, Cell, 2018 (https://doi.org/10.1016/j.cell.2018.05.024). It uses a split-pool technique to capture multi-way chromatin interactions. As in Hi-C experiments, the genomic DNA is cross-linked and fragmented. However, the enzymatic ligation step present in Hi-C experiments is replaced with the ligation of combinatorial barcodes at the ends of the cross-linked DNA fragments. After sequencing, the barcodes can be used to identify clusters of interacting genomic loci. A version of the protocol is also available online from the Guttman lab at http://www.lncrna-test.caltech.edu/protocols/SPRITE_Protocol_July_2018.pdf. The protocol states that the "terminal tags" contain unique sequences of 9 bases. However, such unique sequences appear to vary from 9 to 12 bases in the full list of 96 terminal tag sequences provided by Table S5 of the above article.

My question is: why does the length of unique sequences within the "terminal tags" vary from 9 to 12 bases? In particular, why not use a constant length of 12 bases for all such sequences? A constant length would simplify the data analysis, and using the longer sequences would increase robustness to sequencing errors.

sequence

1 answer

it turns out that varying the length of the unique sequences in the terminal tags helps to avoid libraries of low diversity, which would yield low-quality reads on Illumina sequencers, as explained in Best practices for low diversity sequencing on the NextSeq and MiniSeq systems.

Log in to answer this question.