3 ms·
That is incorrect for scientific data. The limitations are: a) Massive data volumes (~100 Gb - 1 Pb/project) ai) This means that data is typically stored on
by JBorrow 3y ago
That is incorrect for scientific data. The limitations are:
a) Massive data volumes (~100 Gb - 1 Pb/project)
ai) This means that data is typically stored on limited access machines like HPC clusters
bi) This also means that shipping this data around is financially expensive, and cannot be supported purely by small client machines
b) A low number of seeders; scientific data is not exactly popular, and there may be network restrictions on uploads through the typically used networks;
c) The requirement for a data legacy; torrents are fantastic for ephemeral data (e.g. operating system builds), but are terrible for data that must be archived and kept for potentially decades to centuries.
- chaxor 3y agoI'm not sure I agree with this view. If the NIH or DoD, etc serves data over FTP, they can certainly also serve it via IPFS. It would be much better for them to do so. Then, if anyone pulls that data (even if it's only a small dataset, say 2 TB, to their house or lab) it is now more available (since now NIH and other labs are sharing it). I would imagine that the data centers serving these files would be happy to reduce the bandwidth they have to serve by allow other labs to help out as well. It only adds. I don't understand how it subtracts.
- 0cf8612b2e1e 3y agoMost scientific datasets are not that large. For every CERN type study there are 1000x biology papers with n=3 where the collected results sit in a single tab of an Excel document.
- jrumbut 3y agoVery true, but if you're planning on expanding the work of a small, pilot study like that and you don't have the people who were involved in the original you probably need to recreate the study (to shake out the kinks in the protocol, confirm results for yourself, etc). It would be challenging to find a solution robust enough for CERN type data but also simple enough for an n=3 undergraduate research project (that may have yielded some interesting results). I don't know what the solution is there. My intuition is that university libraries could be involved, and that a data librarian could help you get your small study into shape or be embedded at a percentage effort on a large study.
- robwwilliams 3y agoThose do not belong in IPFS. They won’t be replicated and may die.
- jhbadger 3y agoDepends on what you mean by "biology". Maybe some traditional naturalist-style data about where butterfly species were spotted or something like that could just be a spreadsheet, but in more modern biology studies things like RNA-Seq are used, which generate gigabytes or even terabytes of data per paper.
- ISL 3y agoThose biologists might have huge image archives, even if they use microscopes.
- jkh1 3y agoIn biology we now routinely produce datasets in the multiple terrabytes range. It can easily be n = 3 x 10 TB such as for example imaging 3 fly embryos by light sheet microscopy.
- staunton 3y agoScientific datasets like that can be very easily hosted at one of the repositories such as zotero. The only reasons people don't do that is a vague sense of insecurity about having someone declare their analysis botched, vague legal worries, vague unwillingness to do the very small amount of work required to publish data, or the hope to milk a dataset for more papers before anyone else gets a chance.