This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Forum: How to calculate Tajima's D and Fay & Wu's H for unphased data?

Hi,

I have a small number of samples (~10) for my species of interest (non-model organism), so it's almost impossible to phase the data. I am interested in doing some site-frequency spectrum methods to detect positive selection in the genome, but they require the calculation of nucleotide diversity (pi). Is it possible to do so without phasing the data?

Thanks in advance!

evolution genomics nucleotide-diversity genetics

1 answer

Maybe you could use VariScan. However, I don't know if it's the best way to do it for unphased data since you will have to produce 2 sequences for each individual, and therefore randomly assign each variant to one sequence.

Thanks!

Do you know any literature doing the random assignment of variants if the data is unphased?

Just an update. I found a study using your method. They call this process 'haploidize data'.

Log in to answer this question.