I have to analyze a large number of Affymetrix microarray datasets, something I haven't done in many years. I need to select the most informative/appropriate probeset to do DE analysis on. I understand this might not be optimal, but it's needed here.
After reading up I'm confused about the _x_at probesets. Why are they included if not guaranteed to be uniquely mapping to one gene? Is it standard practice to exclude them? Would a reasonable selection process be:
- Exclude _x_at probesets
- Exclude probesets with > 50% of samples in all tested groups (always just case vs control) having absent calls
- Use the probeset with highest median signal in any of the groups being compared
I'm not interested in transcripts here, only gene-level. Would it be better to just choose _at probesets instead of the last step? Or to take the one with the smallest CV within each group?
Had forgotten what a pain MA is :)
EDIT: Some more info. I only have RMA processed data and the probesets have already been mapped to gene and is something I can't change. It's purely the ranking of probesets I'm interested in. Or pooling if that is considered better. I understand this has been asked a million times, but most threads are old and I wasn't able to find something definitive.
0 answers
No answers yet.
Log in to answer this question.