Hi Kevin,
I have used the filtering criteria as per your suggestions, but what i have recently noticed is that with some highly expressed lncRNA genes, particularly with intronic ones, they tend to overlap with highly expressed protein coding genes. When i look at the raw counts generated using lncRNA only annotation, these genes show very high read counts and they pass the FDR and log2FC filters. But when I look at the read counts for the same lncRNA genes generated using the full annotation (whole human genome with lncipedia annotation) the counts for these genes across most samples are down to zero in many cases. I know this probably due to the way htseq regards these counts as ambiguous (i used the default UNION mode in htseq), but in this context, would the use of lncRNA only annotation simply produce an overestimation of the number of long non-coding RNA differentially expressed.
So my question is, if i do use lncRNA only annotation to generate the count matrix with HTSeq and do subsequent DE with DESeq2, is there a particular way to assess if the identified differentially expressed lncRNA genes are indeed being expressed and if so how to filter out these false-positives. I have performed kallisto analysis on the same data with full annotation, but i am having difficulty in understanding how to interpret both pipelines together (kallisto and hisat2/htseq) (Hope that made sense....)
Thanks,
Conor