This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Does anyone know how to transform the following transcript ID's from TCGA isoform data

I'm looking for a way to transform the following transcript ID from TCGA isoform data into ensemble IDs. Examples:

uc011lsn.1
uc010unu.1
uc010uoa.1
uc002bgz.2
uc002bic.2
uc010zzl.1
uc001jiu.2
uc010qhg.1
uc011krn.1
uc003wfr.3
uc003wft.3
uc003wfu.2
uc011kup.1
uc011mlh.1
uc010nib.1
uc010ihw.1
uc004dpj.2
uc010zub.1
uc001qoa.2
uc010ewg.2
uc010ewh.1
uc011cjl.1
uc010mpu.1
uc003ydl.1
uc010mgi.2
...
ucsc isoform tcga

Thank you so much for the response. I've tried this and got gene symbols as well and was wondering now if there is a way to turn it into a ensemble transcript ID such as ENST#########.# or ncbi refseq.

I have a feeling that you're getting your data from a very old database or of some outdated version. These transcripts IDs are not used anymore and haven't been for a long time. Can you tell us where you got the data ? Is it a TCGA RNA-seq expression matrix?

1 answer

You are using the 10+ year old hg19 TCGA data. Why not go to GDC and download the latest data that comes with ensemble ids

Log in to answer this question.