I am Hari Prakash, Pursuing Integrated M.Tech Bioinformatics in Bharathidasan University,Trichy,India.
Currently I was working on Breast cancer project in Micro Array Technology. I have checked and downloaded a dataset from the following link: http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE36693
I have a few questions, namely:
- What does ' +' (plus) and '-'(minus) values indicates in samples column
- how to calculate the pvalue
- why does sample value has highest deviation ex:- 25.6 to 557
Could anyone help me with these?
1 answer
There is no "samples" column in these data. However, if you are referring to the "VALUE" column in individual GSM records, the values are normalized intensity values. The negative values are potentially because the intensity value was below background or because a log transformation was applied. As for the p-value, that is likely calculated by the software used to scan the arrays and is meant as a "detection" p-value and has nothing to do with a gene/probe being differentially-expressed. I say "likely" in this paragraph because the details regarding the array data processing are not clearly specified. You'll need to email the authors or refer to the manuscript to get more details.
As for the "deviation" that you mention in #3, I'm not sure what you mean, so you may need to clarify if you want an answer.
Log in to answer this question.
Regarding 1., which of the files are you looking at?
BTW, I've removed the various instances of "sir" from your post. I realize that you're trying to be polite by doing that, but it ends up excluding women (we already have enough problems with gender-imbalance in this field).
Good job on phrasing the question well, Hari. This time, you have given us your context, the problem you are facing and what you're looking to get out of this discussion. This is good progress. Welcome to the world of scientific forums :)