This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Analysing Microarray file

Hello all, I am trying to analyse a microarray dataset GSE45233 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE45233) but i can see two files here. According to my knowledge i will have to use the "GSE45233_non-normalized.txt" this file. But when i am opening this file it has columns of Avg signal and Detection p-values for each gene. Even the genes are repeated. So how will i find the fold change and p-value of the comparison of each gene.

I am pasting the column names of the file

ID_REF  SYMBOL  6987 - OW2.AVG_Signal   6987 - OW2.Detection Pval   6992 - YOA.AVG_Signal   6992 - YOA.Detection Pval   6982 - YL1.AVG_Signal   6982 - YL1.Detection Pval   6990 - OO2.AVG_Signal   6990 - OO2.Detection Pval   6984 - YW2.AVG_Signal   6984 - YW2.Detection Pval   6986 - OL2.AVG_Signal   6986 - OL2.Detection Pval   6991 - YLC.AVG_Signal   6991 - YLC.Detection Pval   6988 - OW3.AVG_Signal   6988 - OW3.Detection Pval   6993 - YOC.AVG_Signal   6993 - YOC.Detection Pval   6985 - OL1.AVG_Signal   6985 - OL1.Detection Pval   6983 - YW1.AVG_Signal   6983 - YW1.Detection Pval   6989 - OO1.AVG_Signal   6989 - OO1.Detection Pval                                       

What pipeline i should use? I had dealt with RNA-Seq data previously but never with micro-array data. Any help is appreciated.

Thank you

degs microarray

1 answer

The data is Illumina BeadChip microarray. You can read and analyse it using the limma package.

> library(limma)
> y.raw <- read.ilmn("GSE45233_non-normalized.txt.gz",probeid="ID_REF")
Reading file GSE45233_non-normalized.txt.gz ... ...
> y.norm <- neqc(y.raw)
Note: inferring mean and variance of negative control probe intensities from the detection p-values.

This creates an EList object containing background corrected, log2-transformed and normalized intensities.

You can get the sample information from the series_matrix file:

> SampleInfo <- sampleInfoFromGEO("GSE45233_series_matrix.txt.gz")
> SampleInfo
$SampleInfo
   title        geo_accession source_name_ch1                      
1  "6987 - OW2" "GSM1099551"  "Knee meniscus, old, no chondrosis"  
2  "6992 - YOA" "GSM1099552"  "Knee meniscus, young, no chondrosis"
3  "6982 - YL1" "GSM1099553"  "Knee meniscus, young, no chondrosis"
4  "6990 - OO2" "GSM1099554"  "Knee meniscus, old, chondrosis"     
5  "6984 - YW2" "GSM1099555"  "Knee meniscus, young, chondrosis"   
6  "6986 - OL2" "GSM1099556"  "Knee meniscus, old, no chondrosis"  
7  "6991 - YLC" "GSM1099557"  "Knee meniscus, young, no chondrosis"
8  "6988 - OW3" "GSM1099558"  "Knee meniscus, old, chondrosis"     
9  "6993 - YOC" "GSM1099559"  "Knee meniscus, young, chondrosis"   
10 "6985 - OL1" "GSM1099560"  "Knee meniscus, old, no chondrosis"  
11 "6983 - YW1" "GSM1099561"  "Knee meniscus, young, no chondrosis"
12 "6989 - OO1" "GSM1099562"  "Knee meniscus, old, chondrosis"     

$CharacteristicsCh1
   age                 Sex      body mass index (bmi) chondrosis status
1  "old (>40 years)"   "male"   "overweight"          "no chondrosis"  
2  "young (<40 years)" "male"   "obese"               "no chondrosis"  
3  "young (<40 years)" "female" "lean"                "no chondrosis"  
4  "old (>40 years)"   "male"   "obese"               "chondrosis"     
5  "young (<40 years)" "male"   "overweight"          "chondrosis"     
6  "old (>40 years)"   "male"   "lean"                "no chondrosis"  
7  "young (<40 years)" "male"   "lean"                "no chondrosis"  
8  "old (>40 years)"   "male"   "overweight"          "chondrosis"     
9  "young (<40 years)" "female" "obese"               "chondrosis"     
10 "old (>40 years)"   "male"   "lean"                "no chondrosis"  
11 "young (<40 years)" "male"   "overweight"          "no chondrosis"  
12 "old (>40 years)"   "female" "obese"               "chondrosis"     

$CharacteristicsCh2

Downstream analysis follows the usual limma package pipelines.

A comment on sample sizes

If I might add a caution about this dataset. The sample sizes in this dataset are far too small to give conclusions about human patients with arthritis disease with any reasonable confidence. The sample sizes might be ok for lab animals with a clear cut treatment, but are not at all adequate in my opinion to assess differences between human patients with or without a subtle disease. (I would want at least a hundred patients rather than just 12 in order to have worthwhile statistical power.) I would be surprised if a proper limma analysis of this dataset would result in any statistically signifcant DE genes between chondrosis states.

The authors of the original paper simply presented anova p-values without multiple testing for this dataset, so the results that they present are quite likely to have just been due to chance.

Log in to answer this question.