This is a test version of Biostars. For the public version, visit https://www.biostars.org.
How to Interpret PHEWAS hits that are very rare but very significant?

Hello,

I am investigating a particular gene under the hypothesis that this locus associates with a particular disease. On the Global Biobank Engine (GBE) I have come across two missense variants which are ultra-low frequency (GBE reports MAF on the order of e-6; gnomAD has it on the order of e-5) but which seem to associate very strongly with many phenotypes associated with that disease. In particular, for one of these variants there are 7 separate phenotypes which relate to the disease (one of which is the disease itself, several quite specific symptoms, and several less specific) at the p=e-8 level and up (including one at the p=e-10).

As I am new to this sort of analysis, I was hoping for some insight into how to assess the strength of the evidence based on a result like this? If this variant were relatively common I would understand that this signal is very strong, but if it is this rare it must be harbored by just a handful of people. (and it seems unlikely that a handful of people could have all these phenotypes so that confuses me in itself how all of these could pop up). Given the rarity of the allele I am wondering if I can trust the p-values as much as I could otherwise?

I tried to read up on how GBE deals with rare variants. From their FAQ page (https://biobankengine.stanford.edu/faq#aggregate) it seems like they aggregate information across rare variants to try to boost the statistical signal. But the model they cite seems to be a Bayesian one that computes a Bayes factor, so I don't quite understand where the p-value comes from in the first place. Would I be right to think that in my case this particular aggregation was not performed, and the p-value would be that which was output by plink?

Thanks in advance.

frequency phenotype allele phewas p-value

0 answers

No answers yet.

Log in to answer this question.