This is a test version of Biostars. For the public version, visit https://www.biostars.org.
How retaining multi-mapping reads affects CHiCAGO interaction calls (PCHI-C data)

Hi all,

I’m working with promoter capture Hi-C data and I use HiCUP followed by CHiCAGO to call significant interactions. Because my samples come from spermatogenesis, I need to retain multi-mapping reads, which are particularly abundant on the Y chromosome.

To address this, I modified HiCUP as following: for each multi-mapping read with identical alignment scores, I keep one random alignment. Then, I ran CHiCAGO exactly as before, using the standard score threshold of 5.

However, with this multi-mapping retention strategy, only ~50% of the significant interactions identified in the original pipeline (without keeping multimappers) are recovered. I was expecting that most of the interactions would overlap, so this large difference surprised me.

Do you have insights into how keeping multi-mapping reads could affect the CHiCAGO scoring model?

For context, in my dataset, about 40% of reads are multi-mapped.

Any thoughts or suggestions would be greatly appreciated!

reads pchi-c multi-mapped chicago

0 answers

No answers yet.

Log in to answer this question.