I've been thinking about this recently, for your information, here is what I've found,
The PAM50 was actually first trained on qRT-PCR data, http://ascopubs.org/doi/full/10.1200/JCO.2008.18.1370, but this paper seems to have also used microarray data, together with qRT-PCR data, for clustering with 189 breast tumors across 1,906 “intrinsic” genes
Gene Set Reduction Using Prototype Samples and qRT-PCR A minimized gene set was derived from the prototypic samples using the qRT-PCR data for 161 genes that passed FFPE performance criteria established in Mullins et al.21 Several minimization methods were used, including top “N” t test statistics for each group,22 top cluster index scores,23 and the remaining genes after “shrinkage” of modified t test statistics.24 Cross-validation (random 10% left out in each of 50 cycles) was used to assess the robustness of the minimized gene sets. The “N” t test method was chosen due to having the lowest cross-validation (random 10% left out of each iteration) error. The 50 genes selected and their contribution to distinguishing the different subtypes is provided in Appendix Figure A2 (online only).
The 2015 TCGA Breast cancer paper, did adjustment of the RNA-Seq data first, then applied PAM50, http://www.cell.com/cell/abstract/S0092-8674(15)01195-2, in the SI,
To determine breast cancer intrinsic subtypes based on the PAM50 signature, first,
the TCGA mRNA-seq data were subsampled to match the ER distribution of the
training set used for the PAM50. Second, the entire TCGA 817 data set was adjusted
to the median gene expression calculated for the PAM50 genes determined from the ER balanced subset; intrinsic subtyping was then done as previously described (Cancer Genome Atlas, 2012).
But the 2012 breast cancer paper (Cancer Genome Atlas, 2012) used microarray data, I've found inconsistency between the two (~10%). I contacted Dr Perou, he commented that there are always inconsistencies in subtyping between different platforms, which is surprising to me.