3 ms·
Thank you for your efforts! To be clear, I am not advocating for the removal of any files larger than 30 MiB (or any other arbitrary hard limits). It'd be grea
by liberalgeneral 4y ago
Thank you for your efforts!
To be clear, I am not advocating for the removal of any files larger than 30 MiB (or any other arbitrary hard limits). It'd be great of course to flag large files for further review, but the current software doesn't do a great job at crowdsourcing these kinds of tasks (another one being deduplication) sadly.
Given the very little amount of volunteer-power, I'm suggesting that a "lean edition" of LibGen can still be immensely useful to many people.
- ssivark 4y agoFiles are a very bad unit to elevate in importance, and number of files or file size are really bad proxy metrics, especially without considering the statistical distribution of downloads (leave alone the question of what is more "important"!). Eg: Junk that’s less than the size limit is implicitly being valued over good content that happens to be larger in size. Textbooks & reference books will likewise get filtered out with higher likelihood — and that would screw students in countries where they cannot afford them (which might arguable be a more important audience to some, compared to those downloading comics). Etc. After all this, the most likely human response from people who really depend on this platform would be to slice a big file into volumes under the size limit. Seems to be a horrible UX downgrade in the medium to long term for no other reason than satisfying some arbitrary metric of legibility[1]. Here's a different idea -- might it be worthwhile to convert the larger files to better compressed versions eg. PDF -> DJVU? This would lead to a duplication in the medium term, but if one sees a convincing pattern that users switch to the compressed versions without needing to come back to the larger versions, that would imply that the compressed version works and the larger version could eventually be garbage collected. Thinking in an even more open-ended manner, if this corpus is not growing at a substantial rate, can we just wait out a decade or so of storage improvements before this becomes a non-issue? How long might it take for storage to become 3x, 10x, 30x cheaper? [1]: https://www.ribbonfarm.com/2010/07/26/a-big-little-idea-called-legibility/ https://www.ribbonfarm.com/2010/07/26/a-big-little-idea-call...
- didgetmaster 4y ago> can we just wait out a decade or so of storage improvements before this becomes a non-issue? I'm not sure that there is anything on the horizon which would make duplicate data a 'non-issue'. Capacities are certainly growing, so within a decade we might see 100TB HDDs available and affordable 20TB SSDs. But that does not solve the bandwidth issues. It still takes a long, long time to transfer all the data. The fastest HDD is still under 300MB/s which means it takes a minimum of 20 hours to read all the data off a 20TB HDD. That is if you could somehow get it to read the whole thing at the maximum sustained read speed. SSDs are much faster, but it will always be easier to double the capacity than it is to double the speed.
- fragmede 4y agoThe problem isn't the technology, it's the cost. Given a far larger budget, you wouldn't run the hard drives at anywhere near capacity, in order to gain a read speed advantage by running a ton in parallel. That'll let you read 20 TB in a hour if you can afford it. Put it this way; Netflix is able to do 4k video and that's far more intensive.
- titoCA321 4y agoThere's people that contribute to the LibGen ecosystem but unfortunately it in areas that don't really benefit the community. Users don't need another CLI tool for LibGen, nor does the community need another Bot. Unfortunately that's what folks do, make extensions, CLI tools and bots that benefit next to no one and release all over silly willy with no support.