12 ms·
LibGen's Bloat Problem
- powera 4y agoCuration is hard, particularly for a "community" project. Every file is there for a reason, and much of the time, even if it is a stupid reason, removing it means there is one more person opposed to the concept of "curation".
- gizajob 4y agoOne of my favourite places on the internet too. The thing is, you just search for what you want and spend 10 seconds finding the right book and link. While I'd love to mirror whole archive locally, it would really be superfluous because I can only read a couple of quality books at a time anyway, so building my own small archive of annotated PDFs (philosophy is my drug of choice) is better than having the whole. I think it's actually remarkably free of bloat and cruft considering, but maybe I'm not trawling the same corners as you are. Do kind of wish they'd clear out the mobi and djvu versions and make it unified however.
- napier 4y agoIs there a torrent available that would allow straightforward setup of locally storable and accessible Libgen library? For the storage rich but internet connection reliability poor, something like this would be a godsend.
- mdaniel 4y agoThey have a dedicated page where they offer torrents, so pick one of the currently available hostnames: https://duckduckgo.com/?q=libgen+torrent&ia=web https://duckduckgo.com/?q=libgen+torrent&ia=web Obviously, folks can disagree on the "straightforward" part of your comment given the overwhelming number of files we're discussing
- scott_siskind 4y agoWhy would they clear out djvu? It's one of the best/most efficient storage format for scanned books.
- xdavidliu 4y agodjvu is really quite a marvellous format, but I'm only able to read them on Evince (the default pdf reader that comes with Debian, Fedora, and probably a bunch of other distros). For my macbook I need to download a Djvu reader, and for my ipad, I didn't even bother trying because the experience would likely be much worse than Preview / Ibooks.
- eru 4y agoApparently you can install Evince on MacOS as well. But I haven't tried it there. Evince doesn't come by default with Archlinux (my desktop distribution of choice), but I still install it everywhere.
- nsajko 4y ago> Evince doesn't come by default with Archlinux (my desktop distribution of choice) This doesn't make sense; nothing comes "by default" on Arch, but evince is in the official repos as far as I see.
- eru 4y agoA few things come by default on Arch. See the list at https://archlinux.org/packages/core/any/base/ https://archlinux.org/packages/core/any/base/ (many some of these entries like coreutils expand to more packages). Yes, evince is in the official repos. Just like Chromium and Firefox. Or bash, but not any other shell (as far as I can tell).
- MichaelCollins 4y agoCalibre supports djvu on any platform. Deleting djvu books just because Microsoft and Apple don't see fit to support it by default would be a travesty.
- liberalgeneral 4y ago> While I'd love to mirror whole archive locally, it would really be superfluous because I can only read a couple of quality books at a time anyway, [...] I'd love to agree but as a matter of fact LibGen and Sci-Hub are (forced to be) "pirates" and they are more vulnerable to takedowns than other websites. So while I feel no need to maintain a local copy of Wikipedia, since I'm relatively certain that it'll be alive in the next decade, I cannot say the same about those two with the same certainty (not that I think there are any imminent threats to either, just reasoning a priori).
- BossingAround 4y agoSpeaking of mirroring, is there a way to download one big "several-hundred-GB" blob with the full content of the sites for archival purposes? Surely that would act as a failsafe to your problem.
- charcircuit 4y agoI think it's split into a several different torrents since it's so big.
- jart 4y agoWell when a site claims it's for scientific research articles, and you search for "Game Of Thrones" and find this: https://libgen.is/search.php?req=game+of+thrones&lg_topic=libgen&open=0&view=simple&res=25&phrase=1&column=def https://libgen.is/search.php?req=game+of+thrones&lg_topic=li... Someone's going to prison eventually, like The Pirate Bay founders. It's only a matter of time.
- contingencies 4y agoFirst, SciHub != LibGen. Allied projects that clearly share a support base but not identical. Second, please provide a citation for the assertion that sharing copies of printed fiction erodes sales volume. At this point, one may assume that anything that helps to sell computer games and offline swag is cash-in-bank for content producers. Whether original authors get the same royalties is an interesting question. Third, the former Soviet milieu probably isn't currently in the mood to cooperate with western law enforcement.
- sitkack 4y ago> djvu versions This would be disastrous for preservation. Often the djvu versions have no digital version, the books not in print and the publisher isn't around. The djvu archives are often specifically because some old book, really has and had value to people.
- crazygringo 4y agoYeah, I always convert DVJU to PDF (pretty easy) but it never compresses quite as nicely. DJVU is pretty clever in how it uses a mask layer for more efficient compression, and as far as I know, converting to PDF is always done "dumb" -- flattening the DJVU file into a single image and then encoding that image traditionally in PDF. I wonder if it's possible to create a "lossless" DJVU to PDF converter, or something close to it, if the PDF primitives allow it? I'm not sure if they do, if the "mask layer" can be efficiently reproduced in PDF.
- sitkack 4y agoIf you smoke enough algebra, you could use the DJVU algorithm to implement DJVU in PDF with layers. Or heck you could do it in SVG.
- kragen 4y agoYou can't do this in either PDF or SVG, except using JS, which many PDF and SVG viewers don't support.
- RicoElectrico 4y agoI think more often than not djvu blurs out half-tone which can be too aggressive for e.g. newspaper scans and makes a blurry mess.
- maskros 4y agoIt can be done with relative ease. There is a commercial tool somewhere that does it, because I've run across many PDF files that use a DjVu like structure for scanned books. It won't be perfectly lossless, because the IW44 compression of the color layer will need to be recompressed as JPEG or JPEG2000. The JB2 mask layer can be losslessly recompressed as JBIG2 or CCITT G4 Fax.
- gizajob 4y agoMy comment about djvu was mostly just about my own laziness, because (kill me if you need to) I like using Preview on the Mac for reading and annotating, and it doesn't read them, and once they have to live in a djvu viewer, I tend not to read them or mark them up. Same goes for Adobe Acrobat Reader when I'm on Windows on my University's networked PCs.
- kragen 4y agoI wish they'd clear out the PDF versions and replace them with DjVu versions. DjView is better than any PDF reader I've used, and DjVu files are smaller than scanned PDFs.
- james-redwood 4y agoThat’s not going to work, because outside of e-book enthusiasts, few know what DJVU is and even fewer have the technological skills and will to figure out how to open it. A large part of LibGen’s demographic are university students downloading exorbitantly priced textbooks, and given that, having both a pdf and a djvu available would be ideal.
- kragen 4y agoGNOME-based Linux distributions ship with DjVu support by default, and so do MATE and KDE and most document viewers for Android. But even if you're not using Linux, if you're going to spend 50 hours studying a textbook and you're part of a learning community like a university class, with dozens of people facing the same problem, one of you can spend 0.5 hours figuring out how to install DjView so you can read the textbook. That's a much easier problem to solve than finding out about Library Genesis in the first place, not to mention fixing your legal system so it's legal.
- johndough 4y agoThat graph of file size vs. number of files would be much easier to read if it were logarithmic. I guess OP is using matplotlib. In this case, use plt.loglog instead of plt.plot. Also, consider plt.savefig("chart.svg") instead of png.
- liberalgeneral 4y agoHere is the raw data if you are interested: https://paste.debian.net/hidden/77876d00/ https://paste.debian.net/hidden/77876d00/
- johndough 4y agoThanks. Here is a logarithmic plot as SVG: https://files.catbox.moe/zbf35r.svg https://files.catbox.moe/zbf35r.svg On a second thought, a logarithmic histogram might convey even more information, but that would require all file sizes to recompute the bin sizes.
- boarush 4y agoI don't think OP takes into account that there seem to be multiple editions of the same book which are often required by people to refer to. Not everyone wants the latest edition when the class you're in is using some old edition.
- generationP 4y agoIn practice, it's more often the same file with minor edits such as a PDF table of contents added or page numbers corrected. Say, how many distinct editions of this standard text on elementary algebraic geometry are in the following list? http://libgen.rs/search.php?req=cox+little+o%27shea+ideals&open=0&res=25&view=simple&phrase=1&column=def http://libgen.rs/search.php?req=cox+little+o%27shea+ideals&o... Fun fact: the newest one (the 2018 corrected version of the 2015 fourth edition) is not among them.
- boarush 4y agoI like to think that LibGen also serves as a historical database wherein there is a record that a book of a specific edition had its errors corrected. (Although it would be better if errata could be appended to the same file if possible) Yes, for very minor edits, those files should obviously not exist, but for that there would need to be someone who verifies this, which is such an enormous task that likely no one would take up.
- ZeroGravitas 4y agoI notice they have a place to store the OpenLibrary ID, though I've not seen one filled in as yet. OpenLibrary provides both Work and Edition ids, which helps connect different versions. Their database is not perfect either, but it might make more sense to keep the bibliographic data seperate from the copyright contents anyway. https://openlibrary.org/works/OL1849157W/Ideals_varieties_and_algorithms?edition=ia%3Aidealsvarietiesa0000coxd https://openlibrary.org/works/OL1849157W/Ideals_varieties_an...
- liberalgeneral 4y agoIf you are referring to my duplication comments, sure (but even then I believe there are duplicates of the exact same edition of the same book). Though the filtering by filesize is orthogonal to editions etc. so has nothing to do with that.
- bagrow 4y ago> by filtering any "books" (rather, files) that are larger than 30 MiB we can reduce the total size of the collection from 51.50 TB to 18.91 TB I can see problems with a hard cutoff in file size. A long architectural or graphic design textbook could be much larger than that, for instance.
- mananaysiempre 4y agoWhile it’s a bit of an extreme case, the file for a single 15-page article on Monte Carlo noise in rendering[1] is over 50M (as noise should specifically not be compressed out of the pictures). [1] https://dl.acm.org/doi/10.1145/3414685.3417881 https://dl.acm.org/doi/10.1145/3414685.3417881
- TigeriusKirk 4y agoI was just checking my PDFs over 30M because of this post and was surprised to see the DALL-E 2 paper is 41.9M for 27 pages. Lots of images, of course, it was just surprising to see it clock in around a group of full textbooks.
- elteto 4y agoIf I remember correctly images in PDFs can be stored full res but are then rendered to final size, which more often than not in double column research papers end up being tiny.
- RcouF1uZ4gsC 4y ago> I chose 30 MiB somewhat arbitrarily based on my personal e-book library, thinking "30 MiB ought to be enough for anyone" There are books on art and photography and pathology that have multiple high resolution photographs. I don’t think limiting by file size is a good idea.
- c-fe 4y agoThis is a bit anecdotal, but I did upload a book to libgen. I am am avid user of the site, and during my thesis research I was looking for a specific book and could not find it on there. I did however find it on archive.org. I spent the better half of one afternoon extracting the book from archive.org with some Adobe software, since I had to circumvent some DRM and other things, and all of this was also novel to me. In the end I got a scanned PDF, which had several hundred MB. I managed to reduce it to 47 MB, however further reduction was not easily possible at least not with the means I knew or had at my disposal. I uploaded this version to libgen. I do agree that there may be some large files on there, however I dont agree with removing them. I spent some hours to put this book on there so others who need it can access it within seconds. Removing it because it is too large would void all this effort and require future users to go through a similar process than i did just to browse through the book. Also any book published today is most likely available in some ebook format, which is much smaller in size, so I dont think that the size of libgen will continue to grow at the same pace as it is doing now.
- jtbayly 4y agoAgreed. Deduplication should be the bigger goal, in my opinion.
- DiggyJohnson 4y agoEven then, I wouldn’t want a file with text + illustrations to be considered a dupe of a text-only copy of the same work.
- samatman 4y agoIMHO a process which is lossy should never be described as deduplication. What would work out fairly well for this use case is to group files by similarity, and compress them with an algorithm which can look at all 'editions' of a text. This should mean that storing a PDF with a (perhaps badly, perhaps brilliantly) type-edited version next to it would 'weigh' about as much as the original PDF plus a patch.
- aaron695 4y ago> by filtering any "books" (rather, files) that are larger than 30 MiB we can reduce the total size of the collection from 51.50 TB to 18.91 TB, shaving a whopping 32.59 TB Books greater than 30 MiB are all the textbooks. You are killing the knowledge. Also killing a lot of rare things. If you want to do something amazing and small, OCR them. As an example of greater than 30 meg, I grabbed a short story by Greg Bear the other day not available digitally, it was in a 90 meg copy of a 1983 Analog Science Fiction and Fact Side note de-duping is an incredibly hard project, how will you diff a mobi and a epub and then make a decision? Or a decision between a mobi and a mobi? Books also change with time. Even in the 90's kids books from the 60's had been 'edited' These can be hidden gems to collectors. Cover art also.
- agumonkey 4y agoThere are classes of books that are significantly larger than the rest, like medical / biology books. I don't know if they embed vector based images of the whole body or maybe hundreds of images but it's surprising big they are. Who's in to make some large data gathering about unoptimized books and potentially redudant ones ? or maybe trim pdfs (qpdf can optimize a structure to an extent)
- liberalgeneral 4y agoDatabase dumps are available here if you are interested: http://libgen.rs/dbdumps/ http://libgen.rs/dbdumps/ libgen_compact_* is what you are probably looking for, but they are all SQL dumps so you'll need to import them into MySQL first. :/
- agumonkey 4y agothe dumps are not enough, one has too scan the actual file content to assess the quality are you alone in your analysis or are there groups who try to improve lg ?
- dudehere 4y agoSuch efforts have been made in the past but every time ceased at some point for complexity. A workgroup can be made to tackle it, though.
- Retr0id 4y agoIn an ideal world, every book could be given an "importance" score, for some arbitrary value of importance. For example, how often it is cited. This could be customised on a per-user basis, depending on which subjects and time periods you're interested in. Then you can specify your disk size, and solve the knapsack problem to figure the optimal subset of files that you should store. Edit: Curious to see this being downvoted. Is it really that bad of an idea? Or just off-topic?
- DiggyJohnson 4y agoNot to sound blunt, but answering your question on the downvotes (which you probably didn’t deserve, especially without reply). The concept of an importance score feels very centralized and against the federated / free nature of the site. Towards what end? If the “importance score” impacts curation, I am strongly against it. Not only is it icky, but how is it different than a function of popularity?
- Retr0id 4y agoI'm not suggesting reducing the size of the LibGen collection, I'm thinking along the lines of "I have 2TB of disk space spare, and I want to fill it with as much culturally-relevant information as possible". If the entire collection were availble as a torrent (maybe it already is?), I could select which files I wish to download, and then seed. Those who have 52TB to spare would of course aim to store everything, but most people don't. Just as the proposal in the OP would result in the remaining 32.59 TB of data being less well replicated, my approach has the problem that less "popular" files would be poorly replicated, but you could solve that by also selecting some files at random. (e.g. 1.5TB chosen algorithmically, 0.5TB chosen at random).
- liberalgeneral 4y agoI don't think you've deserved the downvotes, and I don't think it's a bad idea either; indeed some coordination as to how to seed the collection is really needed. For instance phillm.net maintains a dynamically updated list of LibGen and Sci-Hub torrents with less than 3 seeders so that people can pick some at random and start seeding: https://phillm.net/libgen-seeds-needed.php https://phillm.net/libgen-seeds-needed.php
- wishfish 4y agoHas anyone ever stumbled across an executable on LibGen? The article mentioned finding them but I've never seen one. I agree with the other comments that LibGen shouldn't purge the larger books. But, in terms of mirrors, it would be nice to have a slimmed down archive I could torrent. 19 TB would be manageable. And would be nice to have a local copy of most of the books.
- hoppyhoppy2 4y agoI saw a book on antenna design on libgen that originally included a CD with software, and that disk image had been uploaded to the site.
- liberalgeneral 4y ago> Has anyone ever stumbled across an executable on LibGen? The article mentioned finding them but I've never seen one. Here is a list of .exe files in LibGen: https://paste.debian.net/hidden/1c82739a/ https://paste.debian.net/hidden/1c82739a/ And a breakdown of file extensions: https://paste.debian.net/hidden/579e319c/ https://paste.debian.net/hidden/579e319c/ > And would be nice to have a local copy of most of the books. Yes! That was my intention—I wasn't advocating for a purge of content but a leaner and more practical version would be amazing.
- wishfish 4y agoThanks for the lists. I was genuinely curious about the exes. Nice to know where they originate. Interesting that over half of them have titles in Cyrillic. I guess not so many English language textbooks (with included CDs) have been uploaded with the data portion intact.
- macintux 4y ago> Yes! That was my intention—I wasn't advocating for a purge of content but a leaner and more practical version would be amazing. Your piece doesn't make that obvious at all, and given how many people here are misunderstanding that point, you might want to update it.
- rolling_robot 4y agoThe graph should be in logarithmic scale to be readable, actually.
- remram 4y agohttps://news.ycombinator.com/item?id=32540202 https://news.ycombinator.com/item?id=32540202
- Invictus0 4y agoStorage space is not a problem, especially not on the order of terabytes. If you want to download all of libgen on a cheap drive, perhaps limit yourself to epub files only. No one needs all of libgen anyway except archivists and data hoarders.
- liberalgeneral 4y agohttps://news.ycombinator.com/item?id=32540854 https://news.ycombinator.com/item?id=32540854
- Invictus0 4y agoYes, that makes you a data hoarder. Normal people would just use one of the many other methods of getting free books, like legal libraries, googling it on Yandex, torrents, asking a friend, etc. Or just actually pay for a book.
- deleted 4y ago[deleted]
- liberalgeneral 4y agoMy target audience is not normal people though, and I don't mean this in the "edgy" sense. The fact that we are having this discussion is very abnormal to begin with, and I think it's great that there are some deviants from the norm who care about the longevity of such projects. I can imagine many students and researchers hosting a mirror of LibGen for their fellows for example.
- Invictus0 4y agoIn that case, just pay whatever it costs to store the data. With AWS glacier it would cost $50 a month.
- Hizonner 4y agoUm, if the goal is to fit what you can onto a 20TB hard drive at home, then nobody is stopping you from choosing your own subset, as opposed to deleting stuff out of the main archive based on ham-handed criteria...
- ad404b8a372f2b9 4y agoThat's funny, I did the same analysis with sci-hub. Back when there was an organized drive to back it up. I downloaded parts of it and wanted to figure out why it was so heavy, seeing as you'd expect articles to be mostly text and very light. There was a similar distribution of file sizes. My immediate instinct was also to cut off the tail-end, but looking at the larger files I realized it was a whole range of good articles that included high quality graphics that were crucial to the research being presented, not poor compression or useless bloat.
- liberalgeneral 4y agoI think Sci-Hub is the opposite since 1 DOI = 1 PDF in its canonical form (straight from the publisher) so neither duplication nor low-quality is the case.
- dredmorbius 4y agoIt does depend on when the work was published. Pre-digital works scanned in without OCR can be larger in size. That's typically works from the 1980s and before. Given the explosion of scientific publishing, that's likely a small fraction of the archive by work though it may be significant in terms of storage.
- dredmorbius 4y agoIt can be illuminating to look at the size of ePub documents. This is in general an HTML container (and compressed), such that file sizes tend to be quite small. A book-length text (~250 pp or more) might be from 0.3 -- 5 MB, and often at the lower end of the scale. Books with a large number of images or graphics, however, can still bloat to 40-50 MB or even more. Otherwise, generally, text-based PDFs (as opposed to scans) are often in the 2--5 MB range, whilst scans can run 40--400 MB. The largest I'm aware of in my own collection is a copy of Lyell's Geography, sourced from Archive.org. It is of course scans of the original 19th century typography. Beautiful to read, but a bit on the weighty side.
- spiffistan 4y agoI've been dreaming of a book decompiler that would some newfangled AI/ML to produce a perfectly typeset copy of an older book; in the same font or similar, recognizing multiple languages and scripts within the work.
- copperx 4y agoIn the same vein, I would like an e-reader that has TeX or InDesign quality typesetting. I'd settle for Knuth-Plass line breaking with decent justification (and hyphenation). At the very least, make it so that headings do not appear at the bottom of a page. Who thought that was OK?
- mjreacher 4y agoI think one of the problems is the lack of a good open source PDF compressor. We have good open source OCR software like ocrmypdf which I've seen used before, but some of the best compressed books I've seen on libgen used some commercial compressor while the open source ones I've used were generally quite lackluster. This applies double so when people are ripping images from another source, combining them into a PDF then uploading as a high resolution PDF which inevitably ends up being between 70-370 MB. How to deal with duplication is also a very difficult problem because there's loads of reasons why things could be duplicated. Take a textbook, I've seen duplicates which contain either one or several of the following: different editions, different printings (of any particular edition), added bookmarks/table of contents for the file, removed blank white pages, removed front/end cover pages, removed introduction/index/copyright/book information pages, LaTeX'd copies of pre-TeX textbooks, OCR'd, different resolution, other kinds of optimization by software that reduces to wildly different file sizes, different file types (eg .chm, PDFs that are straight conversions from epub/mobi), etc. Some of this can be detected by machines, eg usage of OCR but some of the other things aren't easy at all to detect.
- crazygringo 4y agoWhat commercial compressor/performance are you talking about? AFAIK the best compression you see is monochrome pages encoded in Group4, which for example ImageMagick will do which is open source, and ocrmypdf happily works on top of. Otherwise it's just your choice of using underlying JPG, PNG, or JPEG 2000, and up to you to set your desired lossy compression ratio.
- kragen 4y agoPDF also supports JBIG.
- mjreacher 4y agoThis is a pretty common one I see: https://www.pdf-tools.com/en/products/pdf-optimizer/ https://www.pdf-tools.com/en/products/pdf-optimizer/ When I mean optimized I mean maintaining the page quality too. Obviously you can make the PDF look like crap but that's not very useful.
- _Algernon_ 4y agoMy main issue with libgen is its awful search. Can't search by multiple criteria, shitty fuzzy search, and cant filter by file type.
- liberalgeneral 4y agoZ-Library has been innovating a great deal in that regard. Sadly they are not as open/sharing as LibGen mirrors in giving back to the community (in terms of database dumps, torrents, and source code).
- MichaelCollins 4y agoHave you ever used a card catalogue?
- _Algernon_ 4y agoNo. Your point being?
- MichaelCollins 4y agoIn that case your expectations are understandable. People in your generation are accustomed to finding anything in mere seconds. Not very long ago, if it took you a few minutes to find a book in the catalogue you would count yourself lucky. And if your local library didn't have the book you're looking for, you could spend weeks waiting for the book to arrive from another library in the system. Libgen's search certainly isn't as good as it could be, but it's more than good enough. If you can't bear spending a few minutes searching for a book, can you even claim to want that book in the first place? It's hard for me to even imagine being in such a rush that a few minutes searching in a library is too much to tolerate. But then again, I wasn't raised with the expectations of your generation.
- gmjoe 4y agoHonestly, it's not a big problem. First of all, bloat has nothing to do with file size -- EPUB's are often around 2 MB, typeset PDF's are often 2-10 MB (depending on quantity of illustrations), and scanned PDF's are anywhere from 10 MB (if reduced to black and white) to 100 MB (for colors scans, like where necessary for full-color illustrations). The idea of a 30 MB cutoff does nothing to reduce bloat, it just removes many of the most essential textbooks. :( Also it's very rare to see duplicates of 100 MB PDF's. Second, file duplication is there, but it's not really an unwieldy problem right now. Probably the majority of titles have only a single file, many have 2-5 versions, and a tiny minority have 10+. But they're often useful variants -- different editions (2nd, 3rd, 4th) plus alternate formats like reflowable EBUB vs PDF scan. These are all genuinely useful and need to be kept. Most of the unhelpful duplication I see tends to fall into three categories: 1) There are often 2-3 versions of the identical typeset PDF except with a different resolution for the cover page image. That one baffles me -- zero idea who uploads the extras or why. My best guess is a bot that re-uploads lower-res cover page versions? But it's usually like original 2.7 MB becoming 2.3 MB, not a big difference. Feels very unnecessary to me. 2) People (or a bot?) who seem to take EPUB's and produce PDF versions. I can understand how that could be done in a helpful spirit, but honestly the resulting PDF's are so abysmally ugly that I really think people are better off producing their own PDF's using e.g. Calibre, with their own desired paper size, font, etc. Unless there's no original EPUB/MOBI on the site, PDF conversions of them should be discouraged IMHO 3) A very small number of titles do genuinely have like 5+ seemingly identical EPUB versions. These are usually very popular bestselling books. I'm totally baffled here as to why this happens. It does seem like it would be a nice feature to be able to leave some kind of crowdsourced comments/flags/annotations to help future downloaders figure out which version is best for them (e.g. is this PDF an original typeset, a scan, or a conversion? -- metadata from the uploader is often missing or inaccurate here). But for a site that operates on anoynmity, it seems like this would be too open to abuse/spamming. Being able to delete duplicates opens the door to accidental or malicious deleting of anything. I'd rather live with the "bloat", it's really not an impediment to anything at the moment.
- titoCA321 4y agoWhen you look at movie pirates, there's still uploads of Xvid in 2022. Crap goes in as PDF, mobi, epub, txt and comes out as PDF, mobi, DOCX, txt.
- signaru 4y agoI've experienced scanning personal books and also try to reduce them since I'm also concerned with bloat on my (older) mobile reading devices. Unfortunately, there are reasons I cannot upload those, but the procedures might still be helpful for existing scans. Use ScanTailor to clean them up. If there is no need for color/grayscale, have the output strictly black and white. OCR them with Adobe Acrobat ClearScan (or something else, that is what I have). Convert to black and white DJVU (Djvu-Spec). Dealing with color is another thing, and takes some time. I find that using G'MIC's anisotropic smoothing can help with the ink-jet/half-tone patterns. But it's too time consuming to be used for books.
- pronoiac 4y agoI like ScanTailor! I've used ocrmypdf for the OCR and compression steps. It uses lossless JBIG2 by default, at 2 or 3k per page; I'm curious how that compares to DJVU. (And my mistake, pdf and DJVU are competing container formats.)
- signaru 4y agoIf the PDF is from a scanned source, converting it to DJVU with equivalent DPI typically results to about half the file size (figures can vary depending on the specifics of the PDF source).
- repple 4y agoThis book has a great overview of the origins of library genesis. Shadow Libraries: Access to Knowledge in Global Higher Education https://libgen.is/search.php?req=shadow+libraries https://libgen.is/search.php?req=shadow+libraries
- deleted 4y ago[deleted]
- keepquestioning 4y agoCan we put LibGen on the blockchain?
- FabHK 4y agoIn case that was not a joke: No. LibGen is not a trivial amount of data (a few hard disks full). The blockchain can only handle tiny amounts of data very slowly.
- Synaesthesia 4y ago>"30 MiB ought to be enough for anyone" Sometimes you have eg a history book which has a lot of high quality photos, and then it can be quite large.
- idealmedtech 4y agoI'd love to see that distribution at the end with a log-axis for the file size! Or maybe even log-log, depending. Gives a much better sense of "shape" when working with these sorts of exponential distributions
- cbarrick 4y agoThis is a complete nit, but s/an utopia/a utopia/ Even though "utopia" is spelled starting with a vowel, it is pronounced as /juːˈtoʊpiə/, like "yoo-TOH-pee-ə", with a consonant sound at the start. Since the word starts with a consonant sound, the proper indefinite article is "a".
- kevin_thibedeau 4y agoNow you have to convince the intelligentsia how to use the proper article with "history".
- zozbot234 4y agoWhy do we care so much about "history" anyway? Why not "herstory"?
- dudehere 4y agoWhile it's nice to see people reading, learning, and loving libraries, keep in mind the Library Genesis remnants you are typically using are money hogs covering their profiteering under the original altruistic LG disguise. They don't produce forks and link up everyone to work for their own growth. That's not what LG used to be.
- kragen 4y agoMaybe if the objective is preservation, instead of each person saving an entire copy of libgen locally, people [in a country where this is legal] should save N-of-M shares of it. 51.50 TB in a 5-of-M shares setup would be under 11 TB per share; if M were 16-32 or so, and the community remained sufficiently active to replace the shares held by lapsed participants, it would have a good chance of surviving the next big historical period of book-burnings. A 2TB disk apparently costs about US$40 right now (https://www.mercadolibre.com.ar/disco-duro-interno-western-digital-wd20ezaz-2tb-azul/p/MLA15262764 https://www.mercadolibre.com.ar/disco-duro-interno-western-d... for example) so this is a contribution of about US$200 per participant. Plus the risk of being arrested in the future for possessing forbidden information, of course, but maybe the fact that you can't decrypt it without four other participants would reduce that risk. Of course it's also worthwhile to keep plaintext copies of books you actually read, or might want to read, or want to pretend you actually read. My copy of Kenneth Snelson's Art and Ideas is 41 MB and 174 pages (US$0.0008, 240 kB/page); my copy of Boole's Treatise on the Calculus of Finite Differences is 19 MB and 356 pages (US$0.0004, 53 kB/page); my copy of Kevin Carson's Homebrew Industrial Revolution is 3.8 MB and 399 pages (US$0.00008, 9.5 kB/page). If you were to devote a single terabyte to books for yourself at 10 megabytes per book, you'd still have room for 100'000 books, quite a nice library by any historical standard, even if it's small compared to all of libgen. Perhaps a small group of people [again, in a country where this is legal] participating in a distributed storage system could preserve significant fractions of libgen in such a way even in the face of disaster. I think it's reasonable to weight such preservation efforts toward lighter-weight books, but I also think it's easy to screw that up, for example by keeping only books under 30MB, which throws away all the decent scans of many books, leaving only worthless epub versions. At this point, though, it might be more important to create durable physical artifacts encoding this incomparable treasure; many historical events have obliterated communities of learning, leaving only artifacts. The Cambodian Killing Fields are one recent example, but we can also point to the Spanish Catholic zealots burning the khipu and the Maya codices; the Boxers burning most of the last copy of the Yongle Encyclopedia; Qin Shi Huang's Burning of the Books and Burying of the Scholars; the Christians' prohibition on the Egyptian religion, which brought to a close the millennia-long knowledge of hieroglyphs; and Sulla's conquest of Syracuse, in which Archimedes was killed and his knowledge of the integral calculus was lost until Newton.
- antimony51 4y agoHas any effort been made to remove/remedy "all sorts of binary data"?