Hello all, I am trying to analyse a microarray dataset GSE45233 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE45233) but i can see two files here. According to my knowledge i will have to use the "GSE45233_non-normalized.txt" this file. But when i am opening this file it has columns of Avg signal and Detection p-values for each gene. Even the genes are repeated. So how will i find the fold change and p-value of the comparison of each gene.
I am pasting the column names of the file
ID_REF SYMBOL 6987 - OW2.AVG_Signal 6987 - OW2.Detection Pval 6992 - YOA.AVG_Signal 6992 - YOA.Detection Pval 6982 - YL1.AVG_Signal 6982 - YL1.Detection Pval 6990 - OO2.AVG_Signal 6990 - OO2.Detection Pval 6984 - YW2.AVG_Signal 6984 - YW2.Detection Pval 6986 - OL2.AVG_Signal 6986 - OL2.Detection Pval 6991 - YLC.AVG_Signal 6991 - YLC.Detection Pval 6988 - OW3.AVG_Signal 6988 - OW3.Detection Pval 6993 - YOC.AVG_Signal 6993 - YOC.Detection Pval 6985 - OL1.AVG_Signal 6985 - OL1.Detection Pval 6983 - YW1.AVG_Signal 6983 - YW1.Detection Pval 6989 - OO1.AVG_Signal 6989 - OO1.Detection Pval
What pipeline i should use? I had dealt with RNA-Seq data previously but never with micro-array data. Any help is appreciated.
Thank you
1 answer
The data is Illumina BeadChip microarray. You can read and analyse it using the limma package.
> library(limma)
> y.raw <- read.ilmn("GSE45233_non-normalized.txt.gz",probeid="ID_REF")
Reading file GSE45233_non-normalized.txt.gz ... ...
> y.norm <- neqc(y.raw)
Note: inferring mean and variance of negative control probe intensities from the detection p-values.
This creates an EList object containing background corrected, log2-transformed and normalized intensities.
You can get the sample information from the series_matrix file:
> SampleInfo <- sampleInfoFromGEO("GSE45233_series_matrix.txt.gz")
> SampleInfo
$SampleInfo
title geo_accession source_name_ch1
1 "6987 - OW2" "GSM1099551" "Knee meniscus, old, no chondrosis"
2 "6992 - YOA" "GSM1099552" "Knee meniscus, young, no chondrosis"
3 "6982 - YL1" "GSM1099553" "Knee meniscus, young, no chondrosis"
4 "6990 - OO2" "GSM1099554" "Knee meniscus, old, chondrosis"
5 "6984 - YW2" "GSM1099555" "Knee meniscus, young, chondrosis"
6 "6986 - OL2" "GSM1099556" "Knee meniscus, old, no chondrosis"
7 "6991 - YLC" "GSM1099557" "Knee meniscus, young, no chondrosis"
8 "6988 - OW3" "GSM1099558" "Knee meniscus, old, chondrosis"
9 "6993 - YOC" "GSM1099559" "Knee meniscus, young, chondrosis"
10 "6985 - OL1" "GSM1099560" "Knee meniscus, old, no chondrosis"
11 "6983 - YW1" "GSM1099561" "Knee meniscus, young, no chondrosis"
12 "6989 - OO1" "GSM1099562" "Knee meniscus, old, chondrosis"
$CharacteristicsCh1
age Sex body mass index (bmi) chondrosis status
1 "old (>40 years)" "male" "overweight" "no chondrosis"
2 "young (<40 years)" "male" "obese" "no chondrosis"
3 "young (<40 years)" "female" "lean" "no chondrosis"
4 "old (>40 years)" "male" "obese" "chondrosis"
5 "young (<40 years)" "male" "overweight" "chondrosis"
6 "old (>40 years)" "male" "lean" "no chondrosis"
7 "young (<40 years)" "male" "lean" "no chondrosis"
8 "old (>40 years)" "male" "overweight" "chondrosis"
9 "young (<40 years)" "female" "obese" "chondrosis"
10 "old (>40 years)" "male" "lean" "no chondrosis"
11 "young (<40 years)" "male" "overweight" "no chondrosis"
12 "old (>40 years)" "female" "obese" "chondrosis"
$CharacteristicsCh2
Downstream analysis follows the usual limma package pipelines.
A comment on sample sizes
If I might add a caution about this dataset. The sample sizes in this dataset are far too small to give conclusions about human patients with arthritis disease with any reasonable confidence. The sample sizes might be ok for lab animals with a clear cut treatment, but are not at all adequate in my opinion to assess differences between human patients with or without a subtle disease. (I would want at least a hundred patients rather than just 12 in order to have worthwhile statistical power.) I would be surprised if a proper limma analysis of this dataset would result in any statistically signifcant DE genes between chondrosis states.
The authors of the original paper simply presented anova p-values without multiple testing for this dataset, so the results that they present are quite likely to have just been due to chance.
Log in to answer this question.