I have to start with saying that I am a clinically active physician and recently began doing research part-time. I am only beginning to grasp the basics of genomics and bioinformatics.
I am currently analyzing whole genome sequencing data and have stumbled upon two issues. The data that I have access to has already undergone variant calling and I am mostly dealing with VCF files. I am interested in a handful of genes and I am going to determine whether there are any pathogenic variants. There are two issues that I have come across that I hope that anyone here may help me with.
- There are a few hundreds or so weird-looking variants in each individual, e.g. 1-17054677-AT-A,ATT (GRCh38). I interpret this as this individual is ostensibly heterozygote for one deletion of T in one allele, and an insertion of T in the other allele. This variant is located in a low complexity region (LCR) according to the gnomAD browser. I therefore believe that this is an artifact and that the genetic composition in this locus is unkown in both alleles.
Q1: Does anyone have any other interpretation of this?
Q2: In my understanding, variants located in LCRs are often excluded from WGS data due to high error rate. Can anyone confirm this?
Q3: If variants located in LCRs should be excluded, can this be performed on VCF files? If so, which software would you recommend?
- In a few genes, some of the variants have This means that in these particular genes, there are a lot of variants that score high on pathogenicity prediction softwares. One such variant is e.g. an inframe deletion in the KCNJ12 gene (17-21416338-CGAG-C). At the gnomAD online browser, the allele frequency for this variant is 0.5 in all populations. gnomAD gives the flag "inbreeding coefficient". I have read about inbreeding coefficient, and if I interpret this correctly, this has the implication that this variant is also an artifact.
Q1: Does anyone have any other conclusion than this being an artifact?
Q2: What is the reason for this being erroneous? Do we have a problem with the sequencing or is this not a cause for concern?
Q3: Does anyone have any tips on how to exclude these kinds of variants in our analyses (provided of course that they are artifacts)?
If anybody could shed some light on these questions I would be very grateful.
0 answers
No answers yet.
Log in to answer this question.