This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Counting Read Coverage in a masked dataset

Hey all,

Aligned my reads to my reference genome and then removed all sites from this data that corresponded to known locations of repeats.

I know want to provide some meaningful insight into coverage at given sites, currently I have coverage at every nucleotide in the genome (except those of repeats) and want to report something like the average for every 10,000 bases. However, if I did something simple, my data would be skewed by the absence of values associated with the repeats, e.g. dividing 5000 non-repeat derived bases by 10,000.

I feel like there must be a tool to correct for this sort of thing so thought I'd ask.

Thanks,
Toby

sequencing coverage

0 answers

No answers yet.

Log in to answer this question.