Thanks Devon. Can you think of any automated way to do that? For the novel TSS groups that are adjacent to the gene start I could use a cut-off; but for other TSS more downstream that are also next to annotated TSS, I'm not sure how that would be feasible, even with the annotated files.
I'm also wondering with what confidence one can trust the differential expression analysis then - if reads are redistributed to these TSS that are not actually novel, then the value of the fpkm assigned to this TSS and the adjacent canonical TSS may be miscalculated. Would you agree with that rationalle?