Dear David,
Thank you for your answer! Would you have an idea how to do something similar in R?
Below I paste a piece of a file that I referred to (hopefully it works). What I expect would be a short gene symbol in the gene_short_name (v5) column. Instead there are isoform ids there. This is a gene.fpkm_tracking file from CuffDiff output.
I also thought that gene_id column should probably contain gene ids from ensembl, or at least the isoform ids which are now in gene_short_name column... No idea what could go wrong, but it is annoying for downstream analyses, because all I get are the isoform numbers all the time...
Any suggestions are very welcome :) Monika
tracking_id class_code nearest_ref_id gene_id gene_short_name
XLOC_000001 - - XLOC_000001 ENST00000450305,ENST00000456328,ENST00000515242,ENST00000518655
XLOC_000002 - - XLOC_000002 ENST00000469289,ENST00000473358,ENST00000607096
XLOC_000003 - - XLOC_000003 ENST00000594647,ENST00000606857
XLOC_000004 - - XLOC_000004 ENST00000492842
XLOC_000005 - - XLOC_000005 ENST00000335137
XLOC_000006 - - XLOC_000006 ENST00000442987
XLOC_000007 - - XLOC_000007 ENST00000496488
XLOC_000008 - - XLOC_000008 ENST00000419160,ENST00000423728,ENST00000425496,ENST00000431321,ENST00000431812,ENST00000432964,ENST00000440038,ENST00000440163,ENST00000445840,ENST00000453935,ENST00000455207,ENST00000455464,ENST00000514436,ENST00000599771,ENST00000601486,ENST00000601814,ENST00000608420
