4 ms·
I think the most fascinating aspect of this article is the conflict between preservation and curation. The world of books is no different than the world of cont
by ideonexus 10y ago
I think the most fascinating aspect of this article is the conflict between preservation and curation. The world of books is no different than the world of content online: it's mostly low-quality and not worth the reader's time. The authors' observe that for book preservation, building a curated library of high-quality texts demands a certain amount of exclusivity in who's doing the curation. At the same time, this exclusivity increases the centralization of the repository and makes it more prone to being taken down if hosting copyrighted content. Whereas a distributed library reduces the exclusivity, but also reduces the overall quality of the library because anyone can contribute to it.
They cite Wikipedia as an example of these competing qualities, where user-contributions declined as stricter quality controls were put in place. I've watched this debate over what qualifies as "notable" for Wikipedia rage for years now, as the community tries to strike a balance between hosting an expansive encyclopedia with one that hosts relevant content that isn't watered-down with too much trivia.
I'm curious what others think about this conflict? Will algorithms and machine learning one day curate out the literary gems for us?
- soundwave106 10y agoTo me it depends on the meaning of where they are applying quality. For content to be preserved I would personally hope that the direction is more on the preservation angle, personally.. There is very little cost for storing information these days, so there's no real reason to me not to cast a wide net on exactly what content is archived. There is no single good definition of "quality content" after all; even the works considered "top quality" can shift over time, plus there are people out there who really get into niches that a group of curators seeking "top quality content" might miss. (Some of these niches after all deliberately include kitschy or trashy "low quality" content.) From what I gather from the article, the barrier was more on the technical side of preservation. This is easier to see: the "barrier to entry" is more making sure an e-book isn't fuzzy low-resolution junk, or making sure the e-book has correct title, author, etc.