This solution would not work. The transcripts are assembled by Cufflinks. They might not be present in public databases. It would be useful if Cufflinks could output the sequence.
Hello,
I have a file.bam file with mapped genes to the human genome. I uploaded it to galaxy and performed cuff links. it has produced the three outputs
a) gene expression
b) transcript expression
c) assembled transcriptome.
All I need is the gene sequence , with the gene ID, how can I download that ?
Thanks
3 answers
Use the gffread tool that comes bundled with cufflinks
gffread -g genome_reference.fasta -w transcript_sequences.fasta assembled_transcripts.gtf
Once your get gtf file from cufflinks convert it to bed by using gtf2bed utility from bedops from [here][1] then fetch fasta sequence from bed file using bedtools
fastaFrombed -fi hg19.fa -bed gtf2bed.bed -fo gtf2bed.fa -s
hth
[1]; http://bedops.readthedocs.org/en/latest/content/reference/file-management/conversion/gtf2bed.html
Maybe you can use NCBI. At the very top of the web page, you have the chance to write the gene ID name. In the left part, you can choose whether a nucleotide, gene or protein should be shown
Yoy did not specify where the gene ID comes from. There are different alternatives to NCBI depending upon where the gene ID is being taken
Log in to answer this question.