Thanks a lot. I tried this one but it is limited to a specific chromosome and position. What am interested in is to have the UTRs sequences of a list of human genes that aren't loaded on the same chromosome to be able to count the GC content for each UTR alone. If you can help to figure out how to do this using R I'd appreciate.
UTR sequence extraction
How can I extract the 5' and 3' UTR sequences of a list of gene IDs using BiomaRt R package?
• 2,550 views
•
link
1 answer
Probably the biomaRt user guide covers your question. For example, it has an example showing how to Retrieve all 5’ UTR sequences of all genes that are located on chromosome 3 between the positions 185,514,033 and 185,535,839.
• 0 views
•
link
• 0 views
•
link
(following h.mon) Try:
library("biomaRt")
ensembl <- useMart("ensembl", dataset = "hsapiens_gene_ensembl")
ID <- "ENST00000429510.1"
getSequence(
id = ID,
seqType = "5utr",
type = "ensembl_transcript_id_version",
mart = ensembl
)
getSequence(
id = ID,
seqType = "3utr",
type = "ensembl_transcript_id_version",
mart = ensembl
)
Which returns:
> getSequence(
+ id = ID,
+ seqType = "5utr",
+ type = "ensembl_transcript_id_version",
+ mart = ensembl
+ )
5utr ensembl_transcript_id_version
1 ATTCTTGTGAATGTGACACACGATCTCTCCAGTTTCCAT ENST00000429510.1
> getSequence(
+ id = ID,
+ seqType = "3utr",
+ type = "ensembl_transcript_id_version",
+ mart = ensembl
+ )
3utr
1 GACGCAGAAGAAACATGTCCTTCATTCACCAGGCTGAGCTTTCACAGTGCAGTGGTTGGTACGGGACTAAATGTGAGGCTGATGCTCTACACAAGGAAAAACCTGACCTGCGCACAAACCATCAACTCCTCAGCTTTTGGGAACTTGAATGTGACCAAGAAAACCACCTTCATTGTCCATGGATTCAGGCCAACAGGCTCCCCTCCTGTTTGGATGGATGACTTAGTAAAGGGTTTGCTCTCTGTTGAAGACATGAACGTAGTTGTTGTTGATTGGAATCGAGGAGCTACAACTTTAATATATACCCATGCCTCTAGTAAGACCAGAAAAGTAGCCATGGTCTTGAAGGAATTTATTGACCAGATGTTGGCAG
ensembl_transcript_id_version
1 ENST00000429510.1
Combination of 5.7 and 5.8 with "ensembl_transcript_id_version" as type. :-)
• 0 views
•
link
Log in to answer this question.