Thanks a lot for your input.
I think that the main problem is indeed cryptic duplications, as suggested by liorglic. Similar to the paper he referred to, I see tracks of heterozygousity in my data, i.e. entire transcripts with only heterozygous SNPs. In my view, this suggests that these are duplicated in the genome, but the assembler (yes, I used rnaspades) recovered only one copy and the reads from the paralog map to this one copy, which results in spurious heterozygous SNP calling. Other types of RNA (lncRNA etc as suggested by colindaven) are most likely adding to this issue, although I think paralogs missing in the transcriptome are the bigger issue.
That also means that more stringent filtering is not a great thing to do. While I can probably get rid of some of the spurious SNPs, I will still retain some of them and also lose correct ones.
I think that colindaven is right, when he said "a transcriptome is not sufficient for assessing hets". At least not a transcriptome that was created as quick and dirty as mine. Luckily, we are currently sequencing the genome, so I just have to be a more patient...
hi i am looking forward to map SNPs from transcriptome data. refernce genome is also available. but can you point me to a beginners guide. where to start? and list of tools that i will need? i am a beginner in NGS please
Please do not post new questions as
answersin existing threads. If you do a google search withrnaseq snp site:biostars.orgyou will find existing threads. If they don't provide adequate information then please post a new question.