If most of the data is mapping to the genome of your plant then perhaps no need to be concerned. If a large(r) fraction remains unmapped then take those reads and blast to see what you get.
Also, we sometimes filter for chloroplast contaminants using erne-filter. This might be an option (I am never sure if for RNAseq one should filter for chloroplast or not. For DNAseq it helps a lot).
However, I fully agree with genomax. If most of the reads are mapped to your target, then it is not your business, maybe the organism has some strange property (i.e. a GC poor region?)
BTW: Can you disclose which plant is that? It has really high GC!
I am using [GATK somatic-CNV pipeline][1] for some hg38 human sample (normal-tumor matched). In CollectAllelicCounts interval_list ([here][2] and [here][3] the discussions) I have tried to …
Hi! I am using FastQC from Babraham Bioinformatics to analyze Illumina RNAseq output (fastq.gz). In the "Per sequence quality scores" of my data I have …
[jellyfish.histo file][1]<br> [fastqc file][2] I generated a kmer count file using jellyfish and subsequently a histogram, which when plotted in R gave the attached graph. …
maybe it is chloroplast? Or some other contaminant?
Thank you, might be... I don't know how to be ensured where that come from
If most of the data is mapping to the genome of your plant then perhaps no need to be concerned. If a large(r) fraction remains unmapped then take those reads and blast to see what you get.
Thank you, makes sense, I should look at the mapping rate
Also, we sometimes filter for chloroplast contaminants using erne-filter. This might be an option (I am never sure if for RNAseq one should filter for chloroplast or not. For DNAseq it helps a lot). However, I fully agree with genomax. If most of the reads are mapped to your target, then it is not your business, maybe the organism has some strange property (i.e. a GC poor region?)
BTW: Can you disclose which plant is that? It has really high GC!
Thank you. good reasoning