Hi!
I have a set of peaks (4 replicates) and I want to find out how many of the peaks overlapping are high confidence peaks in all replicates.
The peaks that are present in all 4 replicates account for 9000 peaks, obviously they have different p-values (from macs2).
What I intend to know is that, how many of these 9000 peaks are more significantly called in all 4 replicates than compared to other peaks ( a kind of ranking system, where I can say that the top 7000 peaks out of 9000 have X pvalue or correlation in all 4 replicates).
Any suggestions?????
Thank you
1 answer
Take a look at ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia -- in particular the irreproducible discovery rate (IDR) metric.
You can find some practical info on IDR at https://sites.google.com/site/anshulkundaje/projects/idr and more rigorous treatment of it in the PDFs at http://www.encodestatistics.org/
Log in to answer this question.
I would look for number of 9000 peaks encompassing peak summits from all replicates.