Thank for so much for your systematic replies. Our goal is to learn, so we thought free data, would be interesting to just help us learn. So we will try out those software as a learning exercise in both theoretical concepts and practical analyses.
Usually, for human genome, with only 30X coverage, if it is reference genome guided, and not de novo assembly, as was done here by some sequencing company, then for regions that exhibit presence / absence variation between reference and target genomes, wouldn't there be problems with assembly not reflecting actual sequence? We are thinking this way because of a famous paper that compared an African pan genome to the European reference genome - link. And because, as we mentioned in our OP, the parents are African and Asian - neither are European.
By genome annotation, we mean predicting genes, pseudogenes, transposons etc. Is VCF generated agnostic or even without annotation of the new assembly? Is the VCF file generated based just on mapping, or by comparison of annotated gene sequences at same syntenic loci and which are homologous? Sorry if our jargon is a little confused or confusing, but hopefully we've explained our questions clearly enough. Thanks in advance.
I don't think 30X is good enough coverage to make clinically accurate determinations. Also, anything even remotely accurate needs to be vetted by doctors and clinical genetics counselors, even for simple single-gene disorders, as no genotype is associated with a fixed phenotype to a "set in stone" level. We learn new information every day, and ClinVar doesn't really measure up to a clinically usable database.
The questions you ask above need a team of full time experts to consult and explain, it's not something you can expect from an online forum of volunteers.
Thank you for your response. Is there a scientific consensus about the minimum acceptable fold coverage for sequencing in order to draw clinically related conclusions? And is there an open source database like ClinVar that folks use and prefer over ClinVar? Thanks again.
You could try HGMD (which is manually curated with information taken from publications), which is IMO a tad better than CLINVAR, but I doubt that will make a difference. I'm not sure of the preferred coverage for clinical-level accuracy, but mutation data alone cannot predict too many diseases.
In any case, you may want to restrict yourself to pathogenic entries from CLINVAR - ideally, only those that do not have conflicting evidence, where every piece of evidence points to the mutation being pathogenic.
Thank you, gonna use recommendations from you and JC to learn new concepts, may take us at least a few weeks of learning from tutorials with some small and smple test cases to even start the analysis we envision. At that time, we will post any follow up questions / doubts. Also, we think it may be better for us to start with some data that is higher coverage ~ 100X rather than get stuck with a genome assembly or VCF file that will be a hurdle in us learning these analyses. So if you have any suggestions for such a test genome that is open source for download and use, please share. Thanks again.
You can search SRA for datasets at that level of coverage, but I am not sure if you'll find any clinical grade dataset.