6 ms·
This is a bit anecdotal, but I did upload a book to libgen. I am am avid user of the site, and during my thesis research I was looking for a specific book and c
by c-fe 4y ago
This is a bit anecdotal, but I did upload a book to libgen. I am am avid user of the site, and during my thesis research I was looking for a specific book and could not find it on there. I did however find it on archive.org. I spent the better half of one afternoon extracting the book from archive.org with some Adobe software, since I had to circumvent some DRM and other things, and all of this was also novel to me. In the end I got a scanned PDF, which had several hundred MB. I managed to reduce it to 47 MB, however further reduction was not easily possible at least not with the means I knew or had at my disposal. I uploaded this version to libgen.
I do agree that there may be some large files on there, however I dont agree with removing them. I spent some
hours to put this book on there so others who need it can access it within seconds. Removing it because it is too large would void all this effort and require future users to go through a similar process than i did just to browse through the book.
Also any book published today is most likely available in some ebook format, which is much smaller in size, so I dont think that the size of libgen will continue to grow at the same pace as it is doing now.
- jtbayly 4y agoAgreed. Deduplication should be the bigger goal, in my opinion.
- DiggyJohnson 4y agoEven then, I wouldn’t want a file with text + illustrations to be considered a dupe of a text-only copy of the same work.
- samatman 4y agoIMHO a process which is lossy should never be described as deduplication. What would work out fairly well for this use case is to group files by similarity, and compress them with an algorithm which can look at all 'editions' of a text. This should mean that storing a PDF with a (perhaps badly, perhaps brilliantly) type-edited version next to it would 'weigh' about as much as the original PDF plus a patch.
- duskwuff 4y ago> IMHO a process which is lossy should never be described as deduplication. Depends. There are going to be some cases where files aren't literally duplicates, but the duplicates don't add any value -- for example, MOBI conversions of EPUB files, or multiple versions of an EPUB with different publisher-inserted content (like adding a preview of a sequel, or updating an author's bibliography).
- samatman 4y agoSplitting those into two cases: I think getting rid of format conversions (which can, after all, be performed again) is worthwhile, but isn't deduplication, that's more like pruning. Multiple versions of an EPUB with slightly different content is exactly the case where a compression algorithm with an attention span, and some metadata to work with can, get the multiple copies down enough in size that there's no point in disposing of the unique parts.
- ajsnigrutin 4y agoPlus there are a lot of books, where one version is a high quality scan, but no OCR, and the other is OCRed scan (with a bunch of errors, but searching works 80% of the time) and horrible scan quality. Also, some books included appendices, that are scanned in some versions but not in others, plus large posters, that are shrunk to a4 size in one version, split onto multiple a4 pages in another, and one huge page in a third version. Then there are zips of books, containing 1 pdf + eg. example code, libraries, etc (eg progrmaming books).
- CamperBob2 4y agoHave to be careful there. A jihad against duplication means that poor-quality scans will drive out good ones, or prevent them from ever being created. Especially if you're misguided enough to optimize for minimum file size. I agree with samatman's position below: as long as the format is the slightest bit lossy -- and it always will be -- aggressive deduplication has more downsides than upsides.
- willnonya 4y agoWhile intended to agree the duplicates need to be easily identifiable and preferably filterable by quality for bulk downloads.
- exmadscientist 4y agoDeduplication doesn't have to mean removal. It might be just tagging. It would be very nice to be able to fetch the "best filesize" version of the entire collection, then pull down the "best quality" editions of only a few things I'm particularly interested in.
- signaru 4y agoProbably only safe in cases where the files in question are exactly the same binaries (if binary diffing can be automated somehow).
- liberalgeneral 4y agoThank you for your efforts! To be clear, I am not advocating for the removal of any files larger than 30 MiB (or any other arbitrary hard limits). It'd be great of course to flag large files for further review, but the current software doesn't do a great job at crowdsourcing these kinds of tasks (another one being deduplication) sadly. Given the very little amount of volunteer-power, I'm suggesting that a "lean edition" of LibGen can still be immensely useful to many people.
- ssivark 4y agoFiles are a very bad unit to elevate in importance, and number of files or file size are really bad proxy metrics, especially without considering the statistical distribution of downloads (leave alone the question of what is more "important"!). Eg: Junk that’s less than the size limit is implicitly being valued over good content that happens to be larger in size. Textbooks & reference books will likewise get filtered out with higher likelihood — and that would screw students in countries where they cannot afford them (which might arguable be a more important audience to some, compared to those downloading comics). Etc. After all this, the most likely human response from people who really depend on this platform would be to slice a big file into volumes under the size limit. Seems to be a horrible UX downgrade in the medium to long term for no other reason than satisfying some arbitrary metric of legibility[1]. Here's a different idea -- might it be worthwhile to convert the larger files to better compressed versions eg. PDF -> DJVU? This would lead to a duplication in the medium term, but if one sees a convincing pattern that users switch to the compressed versions without needing to come back to the larger versions, that would imply that the compressed version works and the larger version could eventually be garbage collected. Thinking in an even more open-ended manner, if this corpus is not growing at a substantial rate, can we just wait out a decade or so of storage improvements before this becomes a non-issue? How long might it take for storage to become 3x, 10x, 30x cheaper? [1]: https://www.ribbonfarm.com/2010/07/26/a-big-little-idea-called-legibility/ https://www.ribbonfarm.com/2010/07/26/a-big-little-idea-call...
- didgetmaster 4y ago> can we just wait out a decade or so of storage improvements before this becomes a non-issue? I'm not sure that there is anything on the horizon which would make duplicate data a 'non-issue'. Capacities are certainly growing, so within a decade we might see 100TB HDDs available and affordable 20TB SSDs. But that does not solve the bandwidth issues. It still takes a long, long time to transfer all the data. The fastest HDD is still under 300MB/s which means it takes a minimum of 20 hours to read all the data off a 20TB HDD. That is if you could somehow get it to read the whole thing at the maximum sustained read speed. SSDs are much faster, but it will always be easier to double the capacity than it is to double the speed.
- culi 4y agoI've always wanted to contribute to LibGen. Got me through college and has powered my Wikipedia editing hobby Are there any good guides out there for best practices for minimizing files, scanning books, etc?
- generationP 4y agoThere's a bunch. Here's what I do (for black-and-white text; I'm not sure how to deal with more complex scenarios): Scan with 600dpi resolution. Nevermind that this gives huge output files; you'll compress them to something much smaller at the end, and the better your resolution, the stronger compression you can use without losing readability. While scanning, periodically clean the camera or the scanner screen, to avoid speckles of dirt on the scan. The ideal output formats are TIF and PNG; use them if your scanner allows. PDF is also fine (you'll then have to extract the pages into TIF using pdfimages or using ScanKromsator). Use JPG only as a last resort, if nothing else works. Once you have TIF, PNG or JPG files, put them into a folder. Make sure that the files are sorted correctly: IIRC, the numbers in their names should match their order (i.e., blob030 must be an earlier page than blah045; it doesn't matter whether the numbers are contiguous or what the non-numerical characters are). (I use the shell command mmv for convenient renaming.) Import this folder into ScanTailor ( https://github.com/4lex4/scantailor-advanced/releases https://github.com/4lex4/scantailor-advanced/releases ), save the project, and run it through all 6 stages. Stage 1 (Fix Orientation): Use the arrow buttons to make sure all text is upright. Use Q and W to move between pages. Stage 2 (Split Pages): You can auto-run this using the |> button, but you should check that the result is correct. It doesn't always detect the page borders correctly. (Again, use Q and W to move between pages.) Stage 3 (Deskew): Auto-run using |>. This is supposed to ensure that all text is correctly rotated. If some text is still skew, you can detect and fix this later. Stage 4 (Select Content): This is about cutting out the margins. This is the most grueling and boring stage of the process. You can auto-run it using |>, but it will often cut off too much and you'll have to painstakingly fix it by hand. Alternatively (and much more quickly), set "Content Box" to "Disable" and manually cut off the most obvious parts without trying to save every single pixel. Don't worry: White space will not inflate the size of the ultimate file; it compresses well. The important thing is to cut off the black/grey parts beyond the pages. In this process, you'll often discover problems with your scan or with previous stages. You can always go back to previous stages to fix them. Stage 5 (Margins): I auto-run this. Stage 6 (Output): This is important to get right. The despeckling algorithm often breaks formulas (e.g., "..."s get misinterpreted as speckles and removed), so I typically uncheck "Despeckle" when scanning anything technical (it's probably fine for fiction). I also tend to uncheck "Savitzki-Golay smoothing" and "Morphological smoothing" for some reason; don't remember why (probably they broke something for me in some case). The "threshold" slider is important: Experiment with it! (Check which value makes a typical page of your book look crisp. Be mindful of pages that are paler or fatter than others. You can set it for each page separately, but most of the time it suffices to find one value for the whole book, except perhaps the cover.) Note the "Apply To..." buttons; they allow you to promote a setting from a single page to the whole book. (Keep in mind that there are two -- the second one is for the despeckling setting.) Now look at the tab on the right of the page. You should see "Output" as the active one, but you can switch to "Fill Zones". This lets you white-out (or black-out) certain regions of the page. This is very useful if you see some speckles (or stupid write-ins, or other imperfections) that need removal. I try not to be perfectionistic: The best way to avoid large speckles is by keeping the scanner clean at the scanning stage; small ones aren't too big a deal; I often avoid this stage unless I know I got something dirty. Some kinds of speckles (particularly those that look like mathematical symbols) can be confusing in a scan. There is also a "Picture Zones" rider for graphics and color; that's beyond my paygrade. Auto-run stage 6 again at the end (even if you think you've done everything -- it needs to recompile the output TIFFs). Now, go to the folder where you have saved your project, and more precisely to its "out/" subfolder. You should see a bunch of .tif files, each one corresponding to a page. Your goal is to collect them into one PDF. I usually do this as follows: tiffcp *.tif ../combined.tif tiff2pdf -o ../combined.pdf ../combined.tif rm -v ../combined.tif Thus you end up with a PDF in the folder in which your project is. Optional: add OCR to it; add bookmarks for chapters and sections; add metadata; correct the page numbering (so that page 1 is actual page 1). I use PDF-XChangeLite for this all; but use whatever tool you know best. At that point, your PDF isn't super-compressed (don't know how to get those), but it's reasonable (about 10MB per 200 pages), and usually the quality is almost professional. Uploading to LibGen... well, I think they've made the UI pretty intuitive these days :) PS. If some of this is out of date or unnecessarily complicated, I'd love to hear!