Hello,
My experiment has 10 total experimental groups. There are 45 potential comparisons to be made, and 90 directional comparisons to be made. I have obtained differentially enriched peak sets for all all of the comparisons of interest and narrowed those down such that the bed files I am looking to analyze contain the putative NFR regions that make up each of the differential peak sets. Some of the comparisons have quite a large number of differential peaks, while others have only a few. My question now becomes 'What transcription factor binding sites are located within the differential peak sets and how might these differences in accessibility contribute to the the biological differences between the groups being compared?'.
From the reading I did, I thought FIMO would be the best tool to accomplish this, because it is not looking for enriched motifs, in the way that HOMER or other motif finding programs might be. Based on the IP and the data processing up to this point, I do not think looking for enriched motifs over background generated from the genome is sensible. Instead it is simply asking what motifs are there above a set statistical threshold and returning a count with location information. Maybe there is a better way ton ask this question within MEME-SUITE like AME or SEA? I will check those out!
As respectfully as possible, I have some disagreement with the background recommendation made by the MEME team. Setting the background in the way that they have suggested (using the statistics of the peaks being input) drastically reduced the likiehood of any CG rich motif being identified and greatly increase the likelihood of any AT containing motif being called, because the sequences are somewhat CG rich. As such, setting the background in this way eliminates the identification of well-conserved binding sites for transcription factors that we would expect to see from an IP such as this, such as CTCF. In some cases, I would also be establishing background off of as few as ~200 bp, which I think also presents an issue.
Really, I am just trying to find the most supported and justifiable way to proceed with motif identification and analysis. I would love to talk about this further and am happy to provide further detail regarding my experiment and data processing up to this point if beneficial.
Thanks