Counting Read Coverage in a masked dataset
Hey all,
Aligned my reads to my reference genome and then removed all sites from this data that corresponded to known locations of repeats.
I know want to provide some meaningful insight into coverage at given sites, currently I have coverage at every nucleotide in the genome (except those of repeats) and want to report something like the average for every 10,000 bases. However, if I did something simple, my data would be skewed by the absence of values associated with the repeats, e.g. dividing 5000 non-repeat derived bases by 10,000.
I feel like there must be a tool to correct for this sort of thing so thought I'd ask.
Thanks,
Toby
• 483 views
•
link
0 answers
No answers yet.
Log in to answer this question.