4 ms·
I can't speak for all research datasets but I can speak for sequencing based datasets such as whole genome sequencing, RNA sequencing etc. These projects can co
by gww 4y ago
I can't speak for all research datasets but I can speak for sequencing based datasets such as whole genome sequencing, RNA sequencing etc. These projects can cost hundreds of thousands to millions of dollars. Academic and industrial research projects tend to do everything they can to avoid sharing this data for as long as possible.
The academic researchers that generate these datasets and their funding sources such as granting agencies and donors want to see the money being used to generate something important. In this case it certainly in the researchers interest to avoid sharing the data for as long as possible to get as many publications as possible. This makes it easier to get more funding.
There is a push by journals now for researchers to include the data in publicly accessible repositories such as the NCBI SRA or European Genome-phenome Archive. These sites archive the data and make it available for researchers. However, in most cases they require a comprehensive data sharing agreement between your institution and theirs. Most of these agreements have extremely demanding requirements that make it very difficult if not impossible for a requesting institutes legal team to agree to. For example, the institute that owns the data may make demands such as "the right to approve any publications or research work created with this data", which entails sending any publication to them prior to submission and giving them veto power. I understand the reasoning for the agreements their intent was to protect private information from being shared publicly while allowing researchers to use the data. But I think the system is being abused now.
On the other hand, it can be very difficult to get past institutes to make data public. For example, sequencing data contains personally identifiable information and getting approval to share the raw data can be difficult. For prospective studies, Human patients need to consent to their data being shared and may not consent to it being shared publicly.
I can't comment on industry as much. However, I have colleagues who work at companies and publish numerous papers on the same proprietary datasets but refuse to share it with anyone. This is particularly challenging in fields such as cancer research where they may publish some fancy/superior new model for disease risk stratification on their own data but without sharing it's impossible for researchers to independently validate it.