This is a test version of Biostars. For the public version, visit https://www.biostars.org.
sample size for pangenome project

Could anyone provide guidance on calculating the sample size for a population-scale pangenome project? The project aims to:

  1. Develop a national genomic reference resource that is representative of the population,
  2. Integrate genomic data with clinical, epidemiological, and lifestyle information,
  3. Establish a robust data governance framework to support biomedical research, clinical diagnosis, and the implementation of precision medicine.
size pangenome sample

I don’t know, but I think it’s reasonable to say there’s probably not a single generic sample-size answer for this. I think it will depend heavily on your project goals.

Anyway, if I were doing this project, I’d start with a big literature review. I recently took a comparative genomics course that covered (at least some of) this, and because of that, my first stop would be the Human Pangenome Reference Consortium (HPRC).

From there, I’d branch out to papers from HPRC groups and collaborators (through the citation “crawl” that comes with this kind of review work) working on whatever subjects you want to touch with your experimental questions: population representation, graph/pangenome construction, clinical reference databases, common versus rare variation, SNPs and larger structural variants, genotype-phenotype analysis, etc.

I don’t know, but I think it’s reasonable to say there’s probably not a single generic sample-size answer for this. I think it will depend heavily on your project goals.

Agreed. It depends on many factors such as allele frequency, the population structure, and the primary goal. For pangenome projects, you may begin with with a define set of genomes and use rarefaction/saturation analyses to assess whether additional sequencing substantially increases novel variant discovery.

Can you pls help with the specific objectives: Objective 1: Build a population genome reference / pangenome To construct a specific reference genome dataset (reference panel / pangenome), stratified by ethnicity and geography, using long-read sequencing to comprehensively characterize the full spectrum of genetic variation — including single-nucleotide variants, structural variants (SVs), and methylation — across a healthy cohort representative of 54 ethnic groups. This population-level reference serves as a reusable foundation for genetic research, clinical interpretation, and precision medicine. Objective 2: Link genomic data with phenotypic and clinical data To establish the capability to integrate genomic data with standardized phenotypic and clinical data, enabling diagnosis and research in high-priority disease areas — cancer, rare diseases, and pharmacogenomics. This clinical stream connects genomes to clinical records under controlled access, supporting variant interpretation, risk stratification, and genotype-guided treatment.

0 answers

No answers yet.

Log in to answer this question.