5 ms·
So it's between 50 and 60 petabytes of data? I've been wondering how it would be possible for a disparate group of tech-oriented people to make a collection li
by cryptarch 9y ago
So it's between 50 and 60 petabytes of data?
I've been wondering how it would be possible for a disparate group of tech-oriented people to make a collection like that. It would only take a 1000 people with 6 terabytes of storage, which doesn't sound impossible to me.
The main issues I see are:
a) How to share access to the data without exposing yourself?
b) How to make the data discoverable and searchable?
c) How do you ascertain survival of the data?
and optionally: d) How to deal with the freeloader problem?
- denimnerd 9y agoIf we use the private torrent site scene as a model all of those things are pretty much solved. These regulatory agencies go around and play whack a mole on them but they tend to live for a long time and have vast archives when they become mature. see the history of the late what.cd for a rundown of what once existed for music. I think that cheap streaming services have kind of killed the peak potential of the music version of these sites though. It's kind of sad though because what.cd had every single release of every single song catalogued. Streaming sites will only give you 1 or a few.
- cryptarch 9y agoWell... I'd say it isn't solved. What.CD went down. Fuck, that still hurts. People now speak of Google Books as a library of Alexandria, but What.CD was the real thing. Google Books was barely available to anyone, ever. That shouldn't be possible the next time. How though? Distributed metadata curation is a problem we haven't worked out well. I know I haven't. Especially if the metadata is stored on and for data stored on a diverse set of platforms, like "not only BT", but also Freenet, HTTP, FTP, IPFS... It just doesn't exist.
- denimnerd 9y agoI guess separating the meta data from the content in such a way the index can't be targeted for copyright violation.
- cryptarch 9y agoThe thing is, you have to index the pieces of metadata or otherwise make them discoverable, and then decide which parts you trust. This would need some kind of signing/trust-distribution scheme, something like namespaces ("YIFY can only approve movie and show releases because they only do movie and show releases"). It would also need way to blacklist malicious metadata (automatic scanners that publish lists of files with viri?). It's very much non-trivial as far as I can see. That's ignoring the copyright issue, which can mostly be ignored if you somehow make it impractical to prosecute the distributors of the metadata. But I think (part of) the metadata will still be subject to copyright lawsuits, in the context of "the right to be forgotten" and fair-use safe harbors.
- erikpukinskis 9y ago> I'd say it isn't solved. What.CD went down. Fuck, that still hurts. Ethereum. Forever.
- triangleman 9y agoExplain please?
- ams6110 9y agoDon't forget the 100Gb fiber connections.
- sandGorgon 9y agoWhy is it 50 petabytes? If we are talking content of books in some kind of markup, assuming a heavy duty 10mb per book...It would be 1 petabyte. Reasonably, I would assume it would be a few hundred GB. What is in the data that's making it so heavy - the original scanned images ?
- cryptarch 9y agoI think it's the scans, yeah, so ~50mb for the scans and ~1mb for the OCR'd version.