This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Why the number of reads in bam generated by GATK haplotype caller are more than the bam generated after GATK baserecalibrator

I am wondering how the number of reads are more in haplotypecaller bam than in baserecalibrator bam though they dont appear to be duplicates?

baserecalibrator bam gatk read haplotypecaller

enter image description here

As per explanation given here https://gatk.broadinstitute.org/hc/en-us/articles/360040096812-HaplotypeCaller#--bam-output , I noticed two categories of reads in the bam generated from GATK HaplotypeCaller. One set of reads start with HC and another set has original read name.

Can Someone help me in better understanding this scenario.

  1. There are some reads (upper segment; lower reads in pink) which do not have insertion or missense. Their readname is different. What exactly do the reads represent?
  2. Not all haplotype (HC) reads have the insertion. Does that mean it is heterozygous variant?

0 answers

No answers yet.

Log in to answer this question.