Checking all these genes manually against the gene list, I don't believe that any of them are in the TCGA dataset? This seems wrong to me, especially as many of the studies I have looked at to choose these key genes use the TCGA dataset, but I could be missing something.
Hello, I am trying to annotate the GENEIDs downloaded from the TCGA-GBM database for further analysis similar to that done in this paper: https://doi.org/10.1002/jcb.29653
So far I have downloaded all of them and have sorted them into a dataframe with column 1 being the GENEID and the rest of the column being counts for each of the respective samples. What I would like to accomplish is to get the gene name in the first column rather than the GENEID. So far I have tried using a few methods which I will go through below, but still have not managed to match more than ~20 genes or so.
i = 0
for(f in files) {
folder <- gsub(" ", "",paste("gdc_download_20220203_134828.113627/",f))
zipped = list.files(gsub(" ","",paste("gdc_download_20220203_134828.113627/",f)))
tryCatch(temp <- read.table((((gsub(" ", "",paste(folder,'/', zipped)))))),
error = function(e) {
print("Oops")
temp = 0}) #temp[,1] <- gsub("\\.*","",temp[,1])
if (i == 0) {
genes <- data.frame(gene = temp[,1],abundance = temp[,2])
i = i + 1
} else {
temp <- data.frame(gene = temp[,1], abundance = temp[,2])
genes <- inner_join(genes, temp[0:60483,], by = 'gene')
}
}
Here is an example of what the original dataset of GENEIDs look like. (I know I have to remove them but if anyone can clarify what the decimals mean I'd be curious as well!)
"ENSG00000000003.13, ENSG00000000005.5, ENSG00000000419.11, ENSG00000000457.12, ENSG00000000460.15, ENSG00000000938.11, ENSG00000000971.14, ENSG00000001036.12, ENSG00000001084.9, ENSG00000001167.13"
Reading other questions on the board, I first removed the decimals (and numbers after using gsub)
rows <- gsub("\\.", "", genes$gene)
This results in:
"ENSG00000000313,ENSG00000000055,ENSG00000041911,ENSG00000045712,ENSG00000046015"
Additionally, all the datasets I looked at have all GENEIDs of length 15, so I removed leading 0s to reduce the length to 15. (I have the same problems whether or not I do this step so I don't believe this is the main issue, but please let me know if this is correct).
for(i in 1:length(rows)) {
if(nchar(rows[i]) == 17) {
rows[i] = sub("00", "", rows[i])
} else if(nchar(rows[i]) == 16) {
#print(nchar(sub("0", "", rows[i])))
rows[i] = sub("0", "", rows[i])
}
}
I then tried using both BioMart and AnnotationDbi on the data shown to no avail for either. I have additionally tried to ensure I use the correct Annotation data set for hg38 which apparently this dataset is derived from. My code to do this is below:
ourCols <- c("SYMBOL", "GENEID")
ourKeys <- rows
# run the query
annot <- AnnotationDbi::select(EnsDb.Hsapiens.v86,
keys=rows,
columns=ourCols,
keytype="GENEID")
SYMBOL GENEID
1 FUBP1 ENSG00000162613
2 HDAC1P2 ENSG00000233012
3 SCARNA9 ENSG00000254911
4 RP11-369C8.1 ENSG00000258616
5 KIR3DL3 ENSG00000274511
6 CAPN15 ENSG00000282214
7 MIR1910 ENSG00000283416
8 MAP2K4 ENSG00000065559
9 PBLD ENSG00000108187
10 CEP350 ENSG00000135837
11 SMAD4 ENSG00000141646
12 RP11-887P2.3 ENSG00000186076
13 RNU1-115P ENSG00000202199
14 RP11-76H14.2 ENSG00000217769
15 AP000619.5 ENSG00000225678
16 RP1-12G14.5 ENSG00000235168
17 HCG18 ENSG00000235727
18 RP5-1065J22.2 ENSG00000237349
19 MIR2053 ENSG00000238399
20 SETP14 ENSG00000240489
21 CNPY2 ENSG00000257727
This is not a sample of the rows returned, this is ALL of the rows returned out of the 60488 geneIDs in the original.
I have encountered the same issue with biomaRt:
mart <- useMart("ENSEMBL_MART_ENSEMBL")
mart <- useDataset("hsapiens_gene_ensembl", mart)
annotLookup <- getBM(
mart=mart,
attributes=c("ensembl_gene_id",
"external_gene_name"),
filter="ensembl_gene_id",
values=rows,
uniqueRows=TRUE)
ensembl_gene_id external_gene_name
1 ENSG00000065559 MAP2K4
2 ENSG00000108187 PBLD
3 ENSG00000135837 CEP350
4 ENSG00000141646 SMAD4
5 ENSG00000162613 FUBP1
6 ENSG00000186076
7 ENSG00000202199 RNU1-115P
8 ENSG00000217769
9 ENSG00000225678
10 ENSG00000233012 HDAC1P2
11 ENSG00000235168
12 ENSG00000235727 HCG18
13 ENSG00000237349
14 ENSG00000238399 MIR2053
15 ENSG00000240489 SETP14
16 ENSG00000254911 SCARNA9
17 ENSG00000257727 CNPY2
18 ENSG00000258616 LINC02303
19 ENSG00000274511 KIR3DL3
20 ENSG00000282214 CAPN15
21 ENSG00000283416 MIR1910
22 ENSG00000288398
Been trying to solve this for around 2 days now, so it would be very helpful if anyone could point me in the right direction or see an obvious mistake! My reason for doing this is to recreate various methods like the one linked above to validate my own data. The genes I actually care about getting the counts for is as follows: so if there is an easy manual way to do it I could do that as well.
"HOXC10,OSMR,SCARA3,SLC39A10,LHX2,MEOX2,SNAI2,ZNF22,PTPRN,RGS14,G6PC3,IGFBP2,TIMP4,DES,RANBP17,CLEC5A,HOXC11,POSTN,BPIFB2,HOXA13,LRRC10,NELL1,SDR16C5,XIRP2"
Thank you for your help!
1 answer
Using your gene names and BioMart:
Gene stable ID Transcript stable ID Gene name
ENSG00000078898 ENST00000170150 BPIFB2
ENSG00000133110 ENST00000379749 POSTN
ENSG00000133110 ENST00000379747 POSTN
ENSG00000133110 ENST00000379743 POSTN
ENSG00000133110 ENST00000379742 POSTN
ENSG00000133110 ENST00000478947 POSTN
ENSG00000133110 ENST00000473823 POSTN
ENSG00000133110 ENST00000474646 POSTN
ENSG00000133110 ENST00000497145 POSTN
ENSG00000133110 ENST00000541179 POSTN
ENSG00000133110 ENST00000541481 POSTN
ENSG00000106511 ENST00000262041 MEOX2
ENSG00000165512 ENST00000298299 ZNF22
ENSG00000165973 ENST00000298925 NELL1
ENSG00000165973 ENST00000357134 NELL1
ENSG00000165973 ENST00000532434 NELL1
ENSG00000165973 ENST00000528046 NELL1
ENSG00000165973 ENST00000527873 NELL1
ENSG00000165973 ENST00000529595 NELL1
ENSG00000165973 ENST00000524738 NELL1
ENSG00000165973 ENST00000528495 NELL1
ENSG00000165973 ENST00000530672 NELL1
ENSG00000165973 ENST00000529218 NELL1
ENSG00000165973 ENST00000534263 NELL1
ENSG00000165973 ENST00000325319 NELL1
ENSG00000180818 ENST00000515593 HOXC10
ENSG00000180818 ENST00000303460 HOXC10
ENSG00000180818 ENST00000511575 HOXC10
ENSG00000180818 ENST00000514415 HOXC10
ENSG00000180818 ENST00000513413 HOXC10
ENSG00000123388 ENST00000546378 HOXC11
ENSG00000123388 ENST00000243082 HOXC11
ENSG00000168077 ENST00000337221 SCARA3
ENSG00000168077 ENST00000301904 SCARA3
ENSG00000145623 ENST00000502536 OSMR
ENSG00000145623 ENST00000274276 OSMR
ENSG00000145623 ENST00000513831 OSMR
ENSG00000145623 ENST00000509237 OSMR
ENSG00000145623 ENST00000508882 OSMR
ENSG00000258227 ENST00000546910 CLEC5A
ENSG00000258227 ENST00000418498 CLEC5A
ENSG00000258227 ENST00000551012 CLEC5A
ENSG00000258227 ENST00000439991 CLEC5A
ENSG00000258227 ENST00000438351 CLEC5A
ENSG00000258227 ENST00000470595 CLEC5A
ENSG00000258227 ENST00000481301 CLEC5A
ENSG00000198812 ENST00000361484 LRRC10
ENSG00000019549 ENST00000020945 SNAI2
ENSG00000019549 ENST00000649776 SNAI2
ENSG00000019549 ENST00000396822 SNAI2
ENSG00000115457 ENST00000434997 IGFBP2
ENSG00000115457 ENST00000233809 IGFBP2
ENSG00000115457 ENST00000490362 IGFBP2
ENSG00000115457 ENST00000456764 IGFBP2
ENSG00000115457 ENST00000436812 IGFBP2
ENSG00000204764 ENST00000443155 RANBP17
ENSG00000204764 ENST00000523189 RANBP17
ENSG00000204764 ENST00000389118 RANBP17
ENSG00000204764 ENST00000519256 RANBP17
ENSG00000204764 ENST00000522533 RANBP17
ENSG00000204764 ENST00000522066 RANBP17
ENSG00000204764 ENST00000519949 RANBP17
ENSG00000204764 ENST00000519130 RANBP17
ENSG00000204764 ENST00000520864 RANBP17
ENSG00000204764 ENST00000519944 RANBP17
ENSG00000204764 ENST00000523727 RANBP17
ENSG00000204764 ENST00000522734 RANBP17
ENSG00000204764 ENST00000520879 RANBP17
ENSG00000204764 ENST00000522380 RANBP17
ENSG00000204764 ENST00000517629 RANBP17
ENSG00000204764 ENST00000521165 RANBP17
ENSG00000204764 ENST00000518492 RANBP17
ENSG00000204764 ENST00000521759 RANBP17
ENSG00000204764 ENST00000520586 RANBP17
ENSG00000204764 ENST00000523665 RANBP17
ENSG00000204764 ENST00000524364 RANBP17
ENSG00000204764 ENST00000521834 RANBP17
ENSG00000196950 ENST00000458054 SLC39A10
ENSG00000196950 ENST00000409086 SLC39A10
ENSG00000196950 ENST00000444421 SLC39A10
ENSG00000196950 ENST00000359634 SLC39A10
ENSG00000196950 ENST00000430412 SLC39A10
ENSG00000196950 ENST00000412905 SLC39A10
ENSG00000196950 ENST00000464301 SLC39A10
ENSG00000196950 ENST00000418005 SLC39A10
ENSG00000196950 ENST00000465851 SLC39A10
ENSG00000106689 ENST00000560961 LHX2
ENSG00000106689 ENST00000373615 LHX2
ENSG00000106689 ENST00000446480 LHX2
ENSG00000106689 ENST00000488674 LHX2
ENSG00000106031 ENST00000518136 HOXA13
ENSG00000106031 ENST00000649031 HOXA13
ENSG00000170786 ENST00000396721 SDR16C5
ENSG00000170786 ENST00000303749 SDR16C5
ENSG00000170786 ENST00000522671 SDR16C5
ENSG00000054356 ENST00000295718 PTPRN
ENSG00000054356 ENST00000409251 PTPRN
ENSG00000054356 ENST00000460801 PTPRN
ENSG00000054356 ENST00000462351 PTPRN
ENSG00000054356 ENST00000443981 PTPRN
ENSG00000054356 ENST00000423636 PTPRN
ENSG00000054356 ENST00000497977 PTPRN
ENSG00000054356 ENST00000486480 PTPRN
ENSG00000054356 ENST00000489650 PTPRN
ENSG00000054356 ENST00000412847 PTPRN
ENSG00000054356 ENST00000446182 PTPRN
ENSG00000054356 ENST00000440552 PTPRN
ENSG00000054356 ENST00000476930 PTPRN
ENSG00000054356 ENST00000442029 PTPRN
ENSG00000054356 ENST00000451506 PTPRN
ENSG00000054356 ENST00000606213 PTPRN
ENSG00000054356 ENST00000468454 PTPRN
ENSG00000054356 ENST00000484986 PTPRN
ENSG00000054356 ENST00000477819 PTPRN
ENSG00000157150 ENST00000287814 TIMP4
ENSG00000141349 ENST00000269097 G6PC3
ENSG00000141349 ENST00000593115 G6PC3
ENSG00000141349 ENST00000588558 G6PC3
ENSG00000141349 ENST00000585361 G6PC3
ENSG00000141349 ENST00000591696 G6PC3
ENSG00000141349 ENST00000585962 G6PC3
ENSG00000141349 ENST00000590639 G6PC3
ENSG00000141349 ENST00000590253 G6PC3
ENSG00000169220 ENST00000408923 RGS14
ENSG00000169220 ENST00000504631 RGS14
ENSG00000169220 ENST00000514713 RGS14
ENSG00000169220 ENST00000511890 RGS14
ENSG00000169220 ENST00000503110 RGS14
ENSG00000169220 ENST00000512000 RGS14
ENSG00000169220 ENST00000512490 RGS14
ENSG00000169220 ENST00000425155 RGS14
ENSG00000169220 ENST00000502731 RGS14
ENSG00000169220 ENST00000514102 RGS14
ENSG00000169220 ENST00000509289 RGS14
ENSG00000169220 ENST00000503044 RGS14
ENSG00000169220 ENST00000506944 RGS14
ENSG00000175084 ENST00000373960 DES
ENSG00000175084 ENST00000477226 DES
ENSG00000175084 ENST00000492726 DES
ENSG00000175084 ENST00000683013 DES
ENSG00000175084 ENST00000483395 DES
ENSG00000163092 ENST00000409043 XIRP2
ENSG00000163092 ENST00000409728 XIRP2
ENSG00000163092 ENST00000409195 XIRP2
ENSG00000163092 ENST00000672716 XIRP2
ENSG00000163092 ENST00000672277 XIRP2
ENSG00000163092 ENST00000409605 XIRP2
ENSG00000163092 ENST00000409273 XIRP2
ENSG00000163092 ENST00000672671 XIRP2
ENSG00000163092 ENST00000628543 XIRP2
This is a glioblastoma dataset? I downloaded a random HTseq count file from this set from GDC portal. I can find counts for your genes of interest.
$ zgrep -w -f id2 eb80bbc4-7186-4872-86df-90629a53ecb1.htseq.counts
ENSG00000019549.7 162
ENSG00000054356.12 1612
ENSG00000078898.6 8
ENSG00000106031.7 2
ENSG00000106511.5 3728
ENSG00000106689.9 1476
ENSG00000115457.8 2446
ENSG00000123388.4 142
ENSG00000133110.13 247
ENSG00000141349.7 11410
ENSG00000145623.11 1262
ENSG00000157150.4 3962
ENSG00000163092.18 5
ENSG00000165512.4 1040
ENSG00000165973.16 72
ENSG00000168077.12 4782
ENSG00000169220.16 326
ENSG00000170786.11 47
ENSG00000175084.10 100
ENSG00000180818.4 135
ENSG00000196950.12 2765
ENSG00000198812.4 2
ENSG00000204764.11 11
ENSG00000258227.5 209
Thank you! You made me go back and double check why it wasn't working and turns out it was this line:
rows <- gsub("\\.", "", genes$gene)
When it should have been
rows <- gsub("\\.", "", genes$gene)
Thank you so much for your help!
Log in to answer this question.
Decimals are version numbers. Those have to be removed when doing searches.
If you only need GeneID