This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Test Constant Probability At Every Position

I have a Genotype data, it has 10000 positions rows and 22 individuals as column , I calculate how many missing value at every position. the summary of that is ( 0 - 3000 ) means 3000 position no missing ( 1 - 300) , ( 23, 1000), (2-19) now, I need to test whether or not the p = constant vs pi/=pj ( P means the probability that value is missing), I am pretty sure it will not be constant, but I need to test that to support my statement.

test

Please be more vebose, "I have data" is not specific, it is indeed about the least specific and most irrelevant statement you can make about data. Your question is not at all clear, but the formulation indicates, it is probably nonsensical; please help us to help you by giving more details.

  • where are the values originally coming from
  • what type of values, I guess it's p-values of some sort, what is pi/pj?
  • why are some data points missing?
  • why is it important whether or not the values are constant, given which definition of 'constant'?
  • define constant, and why would it be relevant, if e.g. p-values are constant (they are not, for obvious reasons)
  • provide an example

To other members, please do not guess up an answer here.

Dear Michael, I edited my question, hopefully it will be understandable. Thanks

It still isn't. Please address the points raised in Michael's first comment.

Ok, thanks for improving the question, it seems like you want to calculate whether the probability of observing a missing genotype is independent of the sample (column). Did I get that correctly? "constant" isn't the right term here because random variables are never constant, otherwise they won't be called variable ;D

Of course, it would be best if you try to address all the points raised in my first comment, so I could stop guessing.

I do not have any evidence that proof it is random or not, I need to know is the missing randomely distrbuted maybe I can proof that from this question, yes Michael, this also needed is it independent or not. But still need to answer that probability of observing a missing genotype is constant or not.

There still seems to be something lost in translation here. Do you want to look and see if the proportion of missing values is different or not between the samples (columns) or probes/genes/whatever (rows)? If the latter, are you interested in their spacial distribution as well (e.g., if you're looking at a tiling array you might want to consider neighboring probes as a group)?

0 answers

No answers yet.

Log in to answer this question.