This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Using tabix to read a tsv.bgz with genomic coordinate in single column<chromosome>:<loc>:<ref>:<alt>

I'm trying to read some of the data for the UK Biobank Imputed GWAS seen here: https://docs.google.com/spreadsheets/d/1kvPoupSzsSFBNSztMzl04xMoSC3Kcx3CrjVf4yBmESU/edit?ts=5b5f17db#gid=227859291

The data comes in a tsv with the first column having the contig:location:ref:alt ... Is there a simple way for tabix to consume and search this data. It seems like it should be relatively simple... or do I need to pipe the data through something like awk?

tabix

1 answer

According to the description, you have a CRHOM name in the second column and a POS value in the third. If this file is sorted you can tabix index the file like this:

$ tabix -s 2 -b 3 -e 3 variants.tsv.bgz

If this is successfully you can use tabix in the usual way.

fin swimmer

Log in to answer this question.