Hi all,
I’m working with promoter capture Hi-C data and I use HiCUP followed by CHiCAGO to call significant interactions. Because my samples come from spermatogenesis, I need to retain multi-mapping reads, which are particularly abundant on the Y chromosome.
To address this, I modified HiCUP as following: for each multi-mapping read with identical alignment scores, I keep one random alignment. Then, I ran CHiCAGO exactly as before, using the standard score threshold of 5.
However, with this multi-mapping retention strategy, only ~50% of the significant interactions identified in the original pipeline (without keeping multimappers) are recovered. I was expecting that most of the interactions would overlap, so this large difference surprised me.
Do you have insights into how keeping multi-mapping reads could affect the CHiCAGO scoring model?
For context, in my dataset, about 40% of reads are multi-mapped.
Any thoughts or suggestions would be greatly appreciated!
0 answers
No answers yet.
Log in to answer this question.