This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Which Of The 2011 Nar Database Submissions Are Fully Accessible?

Out of curiosity I would like to classify the databases in the 2011 NAR Database Issue by accessibility, specifically which ones offer:

  1. A complete download of the data
  2. A web service (REST or SOAP) which allows automated queries from robots
  3. A bookmarkable website that allows links (i.e. the GET protocol) to individual records without necessarily going through a search form

For example,

Database  Complete Download?  Web Service?   Bookmarkable?   If yes, provide example:
COMBREX   No                  No             Yes             http://combrex.bu.edu/DAI?command=SciBay&fun=proteinCluster&pClusterID=419683

But I can't do this alone.

So if you are interested please visit one of the databases and report your findings (or corrections) here. I'll compile the responses into a spreadsheet and report.

Here is a Google Docs Spreadsheet if you wish to edit it directly (it's wide open): https://spreadsheets.google.com/ccc?key=tZQGRMg24BHKgO4vUjYT5TA&hl=en#gid=0

Per Andra's suggestion, I have decided to put this up on Amazon mturk:

https://www.mturk.com/mturk/searchbar?selectedSearchType=hitgroups&searchWords=nar&minReward=0.00&x=0&y=0&=%2Fsearchbar#

web-service

Jeremy wouldn't it be better to create a shared Google-Spreadsheet for this ?

I agree with Pierre. The first step IMHO would be to extract all the URLs from the abstracts (the abstract should have an URL according to the NAR DB issue guidelines) and dump them into a Google Spreadsheet.

Jeremy, I've suggested to create an article for each DB in wikipedia: http://goo.gl/5jUoK . The infobox would contain the information about the web services.

Shocking how many broken links (404) and busted webapps (500) I've encountered already

Thanks Jeremy ! I didn't understand what was exactly Amazon mturk until now :-)

Hey Jeremy - it is a fun question

Did you make the data available?

I was not able to make the MTurk thing work for some reason - i don't even know if that is still a thing.

The spreadsheet link is still active.

2 answers

I am using amazon's mturk (http://www.mturk.com) for these kind of tasks. I am constantly looking for curated databases that contain citations to pubmed. Going through NAR (and Pathguide.org) manually is not doable. In mturk you can ask so called workers to do specific tasks for which they will be rewarded based on the complexity of the questions asked. For a question like this I would pay between a penny and 5ct for each evaluated website. I will run a new mturk task soon. I will see if I can manage to incorporate your question and will report here.

ASPicDB: a database of annotated transcript and protein variants generated by alternative splicing

Download: no

Web service: no

Bookmarkable: no

This one is odd because it is PHP-driven and non-AJAX, so the search results could have just as easily been bookmarkable

Records are accessed post-hoc job style like in NCBI BLAST so I don't think these are permanent: http://t.caspur.it/ASPicDB/newresults.php?organism=human&job=list1507/job1

Log in to answer this question.