This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Assistance in annotating TCGA genes

Hello, I am trying to annotate the GENEIDs downloaded from the TCGA-GBM database for further analysis similar to that done in this paper: https://doi.org/10.1002/jcb.29653

So far I have downloaded all of them and have sorted them into a dataframe with column 1 being the GENEID and the rest of the column being counts for each of the respective samples. What I would like to accomplish is to get the gene name in the first column rather than the GENEID. So far I have tried using a few methods which I will go through below, but still have not managed to match more than ~20 genes or so.

i = 0 
for(f in files) {   
folder <- gsub(" ", "",paste("gdc_download_20220203_134828.113627/",f))   
zipped = list.files(gsub(" ","",paste("gdc_download_20220203_134828.113627/",f))) 

tryCatch(temp <- read.table((((gsub(" ", "",paste(folder,'/', zipped)))))),
          error = function(e) {
             print("Oops")
            temp = 0})   #temp[,1] <- gsub("\\.*","",temp[,1])   
if (i == 0) {
     genes <- data.frame(gene = temp[,1],abundance = temp[,2])
     i = i + 1   
}  else {
     temp <- data.frame(gene = temp[,1], abundance = temp[,2])
     genes <- inner_join(genes, temp[0:60483,], by = 'gene')   
    }

}

Here is an example of what the original dataset of GENEIDs look like. (I know I have to remove them but if anyone can clarify what the decimals mean I'd be curious as well!)

"ENSG00000000003.13, ENSG00000000005.5, ENSG00000000419.11, ENSG00000000457.12, ENSG00000000460.15, ENSG00000000938.11, ENSG00000000971.14, ENSG00000001036.12, ENSG00000001084.9, ENSG00000001167.13"

Reading other questions on the board, I first removed the decimals (and numbers after using gsub)

rows <- gsub("\\.", "", genes$gene)

This results in:

"ENSG00000000313,ENSG00000000055,ENSG00000041911,ENSG00000045712,ENSG00000046015"

Additionally, all the datasets I looked at have all GENEIDs of length 15, so I removed leading 0s to reduce the length to 15. (I have the same problems whether or not I do this step so I don't believe this is the main issue, but please let me know if this is correct).

for(i in 1:length(rows)) {
  if(nchar(rows[i]) == 17) {
    rows[i] = sub("00", "", rows[i])
  } else if(nchar(rows[i]) == 16) {
    #print(nchar(sub("0", "", rows[i])))
    rows[i] = sub("0", "", rows[i])
  }
}

I then tried using both BioMart and AnnotationDbi on the data shown to no avail for either. I have additionally tried to ensure I use the correct Annotation data set for hg38 which apparently this dataset is derived from. My code to do this is below:

ourCols <- c("SYMBOL", "GENEID")
ourKeys <- rows

# run the query
annot <- AnnotationDbi::select(EnsDb.Hsapiens.v86, 
                               keys=rows, 
                               columns=ourCols, 
                               keytype="GENEID")

          SYMBOL          GENEID
1          FUBP1 ENSG00000162613
2        HDAC1P2 ENSG00000233012
3        SCARNA9 ENSG00000254911
4   RP11-369C8.1 ENSG00000258616
5        KIR3DL3 ENSG00000274511
6         CAPN15 ENSG00000282214
7        MIR1910 ENSG00000283416
8         MAP2K4 ENSG00000065559
9           PBLD ENSG00000108187
10        CEP350 ENSG00000135837
11         SMAD4 ENSG00000141646
12  RP11-887P2.3 ENSG00000186076
13     RNU1-115P ENSG00000202199
14  RP11-76H14.2 ENSG00000217769
15    AP000619.5 ENSG00000225678
16   RP1-12G14.5 ENSG00000235168
17         HCG18 ENSG00000235727
18 RP5-1065J22.2 ENSG00000237349
19       MIR2053 ENSG00000238399
20        SETP14 ENSG00000240489
21         CNPY2 ENSG00000257727

This is not a sample of the rows returned, this is ALL of the rows returned out of the 60488 geneIDs in the original.

I have encountered the same issue with biomaRt:

mart <- useMart("ENSEMBL_MART_ENSEMBL")
mart <- useDataset("hsapiens_gene_ensembl", mart)
annotLookup <- getBM(
  mart=mart,
  attributes=c("ensembl_gene_id",
              "external_gene_name"),
  filter="ensembl_gene_id",
  values=rows,
  uniqueRows=TRUE)

   ensembl_gene_id external_gene_name
1  ENSG00000065559             MAP2K4
2  ENSG00000108187               PBLD
3  ENSG00000135837             CEP350
4  ENSG00000141646              SMAD4
5  ENSG00000162613              FUBP1
6  ENSG00000186076                   
7  ENSG00000202199          RNU1-115P
8  ENSG00000217769                   
9  ENSG00000225678                   
10 ENSG00000233012            HDAC1P2
11 ENSG00000235168                   
12 ENSG00000235727              HCG18
13 ENSG00000237349                   
14 ENSG00000238399            MIR2053
15 ENSG00000240489             SETP14
16 ENSG00000254911            SCARNA9
17 ENSG00000257727              CNPY2
18 ENSG00000258616          LINC02303
19 ENSG00000274511            KIR3DL3
20 ENSG00000282214             CAPN15
21 ENSG00000283416            MIR1910
22 ENSG00000288398                   

Been trying to solve this for around 2 days now, so it would be very helpful if anyone could point me in the right direction or see an obvious mistake! My reason for doing this is to recreate various methods like the one linked above to validate my own data. The genes I actually care about getting the counts for is as follows: so if there is an easy manual way to do it I could do that as well.

"HOXC10,OSMR,SCARA3,SLC39A10,LHX2,MEOX2,SNAI2,ZNF22,PTPRN,RGS14,G6PC3,IGFBP2,TIMP4,DES,RANBP17,CLEC5A,HOXC11,POSTN,BPIFB2,HOXA13,LRRC10,NELL1,SDR16C5,XIRP2"

Thank you for your help!

tcga gbm annotation rnaseq

what the decimals mean I'd be curious as well!

Decimals are version numbers. Those have to be removed when doing searches.

If you only need GeneID

Gene stable ID  Gene name
ENSG00000078898     BPIFB2
ENSG00000133110     POSTN
ENSG00000106511     MEOX2
ENSG00000165512     ZNF22
ENSG00000165973     NELL1
ENSG00000180818     HOXC10
ENSG00000123388     HOXC11
ENSG00000168077     SCARA3
ENSG00000145623     OSMR
ENSG00000258227     CLEC5A
ENSG00000198812     LRRC10
ENSG00000019549     SNAI2
ENSG00000115457     IGFBP2
ENSG00000204764     RANBP17
ENSG00000196950     SLC39A10
ENSG00000106689     LHX2
ENSG00000106031     HOXA13
ENSG00000170786     SDR16C5
ENSG00000054356     PTPRN
ENSG00000157150     TIMP4
ENSG00000141349     G6PC3
ENSG00000169220     RGS14
ENSG00000175084     DES
ENSG00000163092     XIRP2

1 answer

Using your gene names and BioMart:

Gene stable ID  Transcript stable ID    Gene name
ENSG00000078898     ENST00000170150     BPIFB2
ENSG00000133110     ENST00000379749     POSTN
ENSG00000133110     ENST00000379747     POSTN
ENSG00000133110     ENST00000379743     POSTN
ENSG00000133110     ENST00000379742     POSTN
ENSG00000133110     ENST00000478947     POSTN
ENSG00000133110     ENST00000473823     POSTN
ENSG00000133110     ENST00000474646     POSTN
ENSG00000133110     ENST00000497145     POSTN
ENSG00000133110     ENST00000541179     POSTN
ENSG00000133110     ENST00000541481     POSTN
ENSG00000106511     ENST00000262041     MEOX2
ENSG00000165512     ENST00000298299     ZNF22
ENSG00000165973     ENST00000298925     NELL1
ENSG00000165973     ENST00000357134     NELL1
ENSG00000165973     ENST00000532434     NELL1
ENSG00000165973     ENST00000528046     NELL1
ENSG00000165973     ENST00000527873     NELL1
ENSG00000165973     ENST00000529595     NELL1
ENSG00000165973     ENST00000524738     NELL1
ENSG00000165973     ENST00000528495     NELL1
ENSG00000165973     ENST00000530672     NELL1
ENSG00000165973     ENST00000529218     NELL1
ENSG00000165973     ENST00000534263     NELL1
ENSG00000165973     ENST00000325319     NELL1
ENSG00000180818     ENST00000515593     HOXC10
ENSG00000180818     ENST00000303460     HOXC10
ENSG00000180818     ENST00000511575     HOXC10
ENSG00000180818     ENST00000514415     HOXC10
ENSG00000180818     ENST00000513413     HOXC10
ENSG00000123388     ENST00000546378     HOXC11
ENSG00000123388     ENST00000243082     HOXC11
ENSG00000168077     ENST00000337221     SCARA3
ENSG00000168077     ENST00000301904     SCARA3
ENSG00000145623     ENST00000502536     OSMR
ENSG00000145623     ENST00000274276     OSMR
ENSG00000145623     ENST00000513831     OSMR
ENSG00000145623     ENST00000509237     OSMR
ENSG00000145623     ENST00000508882     OSMR
ENSG00000258227     ENST00000546910     CLEC5A
ENSG00000258227     ENST00000418498     CLEC5A
ENSG00000258227     ENST00000551012     CLEC5A
ENSG00000258227     ENST00000439991     CLEC5A
ENSG00000258227     ENST00000438351     CLEC5A
ENSG00000258227     ENST00000470595     CLEC5A
ENSG00000258227     ENST00000481301     CLEC5A
ENSG00000198812     ENST00000361484     LRRC10
ENSG00000019549     ENST00000020945     SNAI2
ENSG00000019549     ENST00000649776     SNAI2
ENSG00000019549     ENST00000396822     SNAI2
ENSG00000115457     ENST00000434997     IGFBP2
ENSG00000115457     ENST00000233809     IGFBP2
ENSG00000115457     ENST00000490362     IGFBP2
ENSG00000115457     ENST00000456764     IGFBP2
ENSG00000115457     ENST00000436812     IGFBP2
ENSG00000204764     ENST00000443155     RANBP17
ENSG00000204764     ENST00000523189     RANBP17
ENSG00000204764     ENST00000389118     RANBP17
ENSG00000204764     ENST00000519256     RANBP17
ENSG00000204764     ENST00000522533     RANBP17
ENSG00000204764     ENST00000522066     RANBP17
ENSG00000204764     ENST00000519949     RANBP17
ENSG00000204764     ENST00000519130     RANBP17
ENSG00000204764     ENST00000520864     RANBP17
ENSG00000204764     ENST00000519944     RANBP17
ENSG00000204764     ENST00000523727     RANBP17
ENSG00000204764     ENST00000522734     RANBP17
ENSG00000204764     ENST00000520879     RANBP17
ENSG00000204764     ENST00000522380     RANBP17
ENSG00000204764     ENST00000517629     RANBP17
ENSG00000204764     ENST00000521165     RANBP17
ENSG00000204764     ENST00000518492     RANBP17
ENSG00000204764     ENST00000521759     RANBP17
ENSG00000204764     ENST00000520586     RANBP17
ENSG00000204764     ENST00000523665     RANBP17
ENSG00000204764     ENST00000524364     RANBP17
ENSG00000204764     ENST00000521834     RANBP17
ENSG00000196950     ENST00000458054     SLC39A10
ENSG00000196950     ENST00000409086     SLC39A10
ENSG00000196950     ENST00000444421     SLC39A10
ENSG00000196950     ENST00000359634     SLC39A10
ENSG00000196950     ENST00000430412     SLC39A10
ENSG00000196950     ENST00000412905     SLC39A10
ENSG00000196950     ENST00000464301     SLC39A10
ENSG00000196950     ENST00000418005     SLC39A10
ENSG00000196950     ENST00000465851     SLC39A10
ENSG00000106689     ENST00000560961     LHX2
ENSG00000106689     ENST00000373615     LHX2
ENSG00000106689     ENST00000446480     LHX2
ENSG00000106689     ENST00000488674     LHX2
ENSG00000106031     ENST00000518136     HOXA13
ENSG00000106031     ENST00000649031     HOXA13
ENSG00000170786     ENST00000396721     SDR16C5
ENSG00000170786     ENST00000303749     SDR16C5
ENSG00000170786     ENST00000522671     SDR16C5
ENSG00000054356     ENST00000295718     PTPRN
ENSG00000054356     ENST00000409251     PTPRN
ENSG00000054356     ENST00000460801     PTPRN
ENSG00000054356     ENST00000462351     PTPRN
ENSG00000054356     ENST00000443981     PTPRN
ENSG00000054356     ENST00000423636     PTPRN
ENSG00000054356     ENST00000497977     PTPRN
ENSG00000054356     ENST00000486480     PTPRN
ENSG00000054356     ENST00000489650     PTPRN
ENSG00000054356     ENST00000412847     PTPRN
ENSG00000054356     ENST00000446182     PTPRN
ENSG00000054356     ENST00000440552     PTPRN
ENSG00000054356     ENST00000476930     PTPRN
ENSG00000054356     ENST00000442029     PTPRN
ENSG00000054356     ENST00000451506     PTPRN
ENSG00000054356     ENST00000606213     PTPRN
ENSG00000054356     ENST00000468454     PTPRN
ENSG00000054356     ENST00000484986     PTPRN
ENSG00000054356     ENST00000477819     PTPRN
ENSG00000157150     ENST00000287814     TIMP4
ENSG00000141349     ENST00000269097     G6PC3
ENSG00000141349     ENST00000593115     G6PC3
ENSG00000141349     ENST00000588558     G6PC3
ENSG00000141349     ENST00000585361     G6PC3
ENSG00000141349     ENST00000591696     G6PC3
ENSG00000141349     ENST00000585962     G6PC3
ENSG00000141349     ENST00000590639     G6PC3
ENSG00000141349     ENST00000590253     G6PC3
ENSG00000169220     ENST00000408923     RGS14
ENSG00000169220     ENST00000504631     RGS14
ENSG00000169220     ENST00000514713     RGS14
ENSG00000169220     ENST00000511890     RGS14
ENSG00000169220     ENST00000503110     RGS14
ENSG00000169220     ENST00000512000     RGS14
ENSG00000169220     ENST00000512490     RGS14
ENSG00000169220     ENST00000425155     RGS14
ENSG00000169220     ENST00000502731     RGS14
ENSG00000169220     ENST00000514102     RGS14
ENSG00000169220     ENST00000509289     RGS14
ENSG00000169220     ENST00000503044     RGS14
ENSG00000169220     ENST00000506944     RGS14
ENSG00000175084     ENST00000373960     DES
ENSG00000175084     ENST00000477226     DES
ENSG00000175084     ENST00000492726     DES
ENSG00000175084     ENST00000683013     DES
ENSG00000175084     ENST00000483395     DES
ENSG00000163092     ENST00000409043     XIRP2
ENSG00000163092     ENST00000409728     XIRP2
ENSG00000163092     ENST00000409195     XIRP2
ENSG00000163092     ENST00000672716     XIRP2
ENSG00000163092     ENST00000672277     XIRP2
ENSG00000163092     ENST00000409605     XIRP2
ENSG00000163092     ENST00000409273     XIRP2
ENSG00000163092     ENST00000672671     XIRP2
ENSG00000163092     ENST00000628543     XIRP2

Checking all these genes manually against the gene list, I don't believe that any of them are in the TCGA dataset? This seems wrong to me, especially as many of the studies I have looked at to choose these key genes use the TCGA dataset, but I could be missing something.

This is a glioblastoma dataset? I downloaded a random HTseq count file from this set from GDC portal. I can find counts for your genes of interest.

$ zgrep -w -f id2 eb80bbc4-7186-4872-86df-90629a53ecb1.htseq.counts
ENSG00000019549.7   162
ENSG00000054356.12  1612
ENSG00000078898.6   8
ENSG00000106031.7   2
ENSG00000106511.5   3728
ENSG00000106689.9   1476
ENSG00000115457.8   2446
ENSG00000123388.4   142
ENSG00000133110.13  247
ENSG00000141349.7   11410
ENSG00000145623.11  1262
ENSG00000157150.4   3962
ENSG00000163092.18  5
ENSG00000165512.4   1040
ENSG00000165973.16  72
ENSG00000168077.12  4782
ENSG00000169220.16  326
ENSG00000170786.11  47
ENSG00000175084.10  100
ENSG00000180818.4   135
ENSG00000196950.12  2765
ENSG00000198812.4   2
ENSG00000204764.11  11
ENSG00000258227.5   209

Thank you! You made me go back and double check why it wasn't working and turns out it was this line:

rows <- gsub("\\.", "", genes$gene)

When it should have been

rows <- gsub("\\.", "", genes$gene)

Thank you so much for your help!

Log in to answer this question.