5 ms·
The entirety of Library Genesis (about 2.7 million books, fairly poorly curated) can also be downloaded, its somewhere around 40 TB altogether of significantly
by ralph87 6y ago
The entirety of Library Genesis (about 2.7 million books, fairly poorly curated) can also be downloaded, its somewhere around 40 TB altogether of significantly more recent books.
- GordonS 6y agoIn plain text format though? AFAIK, most is in PDF, EPUB or mobi format. I'd presume it's not too difficult to extract text from the latter two, but extracting text from PDFs is far from simple, and something you get working for 1 PDF won't necessarily work on another.
- heimatau 6y agoIs there a way to submit to it? Also, is there a way to download the entire thing?
- ralph87 6y agoThey release periodic delta torrents of the entire collection. I won't link it here, but it's trivial to find
- gwern 6y agoThe problem with LibGen, and why EleutherAI has been avoiding it, is that most of it is PDFs and most of the PDFs are scans; OCR layers are typically incredibly crummy. Even if you re-OCR them all with Tesseract or something (which will take quite a while, based on how long Tesseract takes to OCR my books even using parallelism on a Threadripper), the OCR will still be awful data. Do you really want to add that to your training dataset...? Far from obvious, and not when there are so many other pools of text like Arxiv which aren't so insoluble.
- ralph87 6y agoIt of course depends entirely on what kind of tool you are trying to build. PDFs capture a ton of visual information lost in plain text. I know the parent article is discussing construction of a text model, but what about say, if you were attempting to build the ultimate AI page reflow tool? etc.