4 ms·
Interesting project, particularly with the choice of IPFS and DCAT -- something I'll have to look into. There have been other efforts to handle mostly file-bas
by DocSavage 8y ago
Interesting project, particularly with the choice of IPFS and DCAT -- something I'll have to look into. There have been other efforts to handle mostly file-based scientific data with versioning in both distributed (Dat https://blog.datproject.org/tag/science/ https://blog.datproject.org/tag/science/) and centralized ways (DataHub https://datahub.csail.mit.edu/www/ https://datahub.csail.mit.edu/www/). Juan Benet visited our research center to give a talk about IPFS a few years ago. Really fantastic stuff.
I'm the creator of DVID (http://dvid.io http://dvid.io), which has an entirely different approach to how we might handle distributed versioning of scientific data primarily at a larger scale (100 GB to petabytes). Like Qri and IPFS, DVID is written in Go. Our research group works in Connectomics. We start with massive 3D brain image volumes and apply automated and manual segmentation to mine the neurons and synapses of all that data. There's also a lot of associated data to manage the production of connectomes.
One of our requirements, though, is having low-latency reads and writes to the data. We decided to create a Science API that shields clients from how the data is actually represented, and for now, have used an ordered key-value stores for the backend. Pluggable "datatypes" provide the Science API and also translate requests into the underlying key-value pairs, which are the units for versioning. It's worked out pretty well for us and I'm now working on overhauling the store interface and improving the movement of versions between servers. At our scale, it's useful to be able to mail a hard drive to a collaborator to establish the base DAG data and then let them eventually do a "pull request" for their relatively small modifications.
We've published some of our data online (http://emdata.janelia.org http://emdata.janelia.org) and visitors can actually browse through the 3d images using a Google-developed web app, Neuroglancer. It's running on a relatively small VM so I imagine any significant HN traffic might crush it :/ We are still figuring out the best way to handle the public-facing side.
I think a lot of people are coming up with their own ideas about how to version scientific data, so maybe we should establish a meeting or workshop to discuss how some of these systems might interoperate? The RDA (https://rd-alliance.org/ https://rd-alliance.org/)
has been trying to establish working groups and standards, although they weren't really looking at distributed versioning a few years ago. We need something like a Github for scientific data where papers can reference data at a particular commit and then offer improvements through pull requests.
- amirouche 8y ago> We need something like a Github for scientific data where papers can reference data at a particular commit and then offer improvements through pull requests. exactly my thought, do you know any working group that is working toward that goal?
- DocSavage 8y agoIf by working group you mean a cross-company collection of people, I don't know of any or I would've joined them :) I've been working toward that goal for the last 5 years, but primarily with an eye to our kinds of data problems in the Connectomics field. I've been meaning to look at RDA again but reluctant to start a working group myself.
- b_fiive 8y agoHey DocSavage! I'm one of these Qri folks, I'd love to see that working group exist. I have a friend or two at the RDA. maybe we should get an email going on the subject? Projects like these are bigger than any one company or tool :)
- DocSavage 8y agoAgreed. Will follow up on email through your Qri contact page.
- b_fiive 8y agodelightful. thanks!
- amirouche 8y agoI started a awesome list dubbed "awesome data distribution" feedback welcome at https://github.com/amirouche/awesome-data-distribution https://github.com/amirouche/awesome-data-distribution
- benhamner 8y ago