This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Using unpublished sequence data from Sequence Read Archive (SRA)

Hello, I am studying the modeling of RNA-Seq count data. Recently, I am doing some kind of meta-analysis for the multiple dataset deposited in SRA database. I came across one dataset that has not been published in the journal (I searched for every possible combinations of multiple identifiers for dataset).

My question is that, is it okay, or ethical to use unpublished sequence data deposited in SRA in our studies and publish the results? My thought is that it is somewhat unethical, so we should contact those who deposited the sequence and get the permission or something, but want to hear the opinion of the community.

Sincerely,

sequence sequencing sra archive

3 answers

My thought is that it is somewhat unethical, so we should contact those who deposited the sequence and get the permission or something

As a courtesy let the depositor know that you intend to use/publish with this data. Give them a time limit to respond and let you know if they have any objection/comments. If you had by mistake made data public you would want someone to do the same.

good point. very political correct behaviour indeed.

if the data is publicly available in SRA, it's free to use I would say, no matter if it has been published or not (who knows, it might never get published?)

It's not like you made it public, no. So I wouldn't feel bad to use such data.

If the data is publicly available, nobody will accuse you of stealing the results. However, I would still check with the authors, because someone put work and money into producing the data. By publishing their data I may be scooping them and affecting their careers and/or livelihoods.

All of us would like to believe that if the unpublished data is out there, it is because the authors are altruistic or have lost interest in data. Both of those explanations suit our purpose as external users of data. The fact is that some datasets are made public before publication because of journal requirements, funding agency or institutional constraints, and probably other reasons we don't want to think about. The authors may have been forced to make the data public, even though they are planning to publish it.

I would send an e-mail to the authors and find out what is going on. Who knows, a fruitful collaboration may come out of it.

Log in to answer this question.