6 ms·
Honestly I don't think it would be that costly, but it would take a pretty long time to put together. I have a (few years old) copy of Library Genesis converted
by bhaney 1y ago
Honestly I don't think it would be that costly, but it would take a pretty long time to put together. I have a (few years old) copy of Library Genesis converted to plaintext and it's around 1TB. I think libgen proper was 50-100TB at the time, so we can probably assume that AA (~1PB) would be around 10-20TB when converted to plaintext. You'd probably spend several weeks torrenting a chunk of the archive, converting everything in it to plaintext, deleting the originals, then repeating with a new chunk until you have plaintext versions of everything in the archive. Then indexing all that for full text search would take even more storage and even more time, but still perfectly doable on commodity hardware.
The main barriers are going to be reliably extracting plaintext from the myriad of formats in the archive, cleaning up the data, and selecting a decent full text search database (god help you if you pick wrong and decide you want to switch and re-index everything later).
- notpushkin 1y agoI think there’s a couple ways to improve it: 1. There’s a lot of variants of the same book. We only need one for the index. Perhaps for each ISBN, select the format easiest to parse. 2. We can download, convert and index top 100K books first, launch with these, and then continue indexing and adding other books.
- palmfacehn 1y agoThere should be a way to leverage compression when storing multiple editions of the same book.
- bawolff 1y agoFrom a good search perspective though you probably dont want 500 different versions of the same book popping up for a query
- palmfacehn 1y agoAgreed. I would prefer to see a single result for a single title. The option of pursuing different editions should follow from there.
- qingcharles 1y agoAnd without some sort of weighting system, it wouldn't even know which one is the best one to show the user.
- notpushkin 1y agoWe’ll also need to consider that some versions might be easier to index even though the user would prefer another version. E.g. if we have a TXT and EPub, we might want to index TXT (if it’s clean enough), but present user with EPub (with formatting and stuff). But it’s not a huge problem actually: just link to the search page instead and let the user decide what they want to download.
- deleted 1y ago[deleted]
- throwup238 1y agoHow are you going to download the top 100k? The only reasonable way to download that many books from AA or Libgen is to use the torrents, which are sorted sequentially by upload date. I tried to automate downloading just a thousand books and it was unbearably slow, from IPFS or the mirrors both. I ended up picking the individual files out of the torrents. Even just identifying or deduping the top 100k would be a significant task.
- notpushkin 1y agoFor each book they store its exact location in the torrent files. You can see on the book page, e.g.: collection “ia” → torrent “annas-archive-ia-acsm-n.tar.torrent” → file “annas-archive-ia-acsm-n.tar” (extract) → file “notesonsynthesis0000unse.pdf” But probably you should get it from the database dumps they provide instead of hammering the website. So you come up with a list of books you want to prioritize, search the DB for torrent name and file to download, download only the files you need, and extract them. You’ll probably end up with quite a few more books, which you may index or skip for now, but it is certainly doable.
- WillAdams 1y agoThe thing is, for an ISBN, that is one edition, by one publisher and one can easily have the same text under 3 different ISBNs from one publisher (hardcover, trade paperback, mass-market paperback). I count 80+ editions of J.R.R. Tolkien's _The Hobbit_ at: https://tolkienlibrary.com/booksbytolkien/hobbit/editions.php https://tolkienlibrary.com/booksbytolkien/hobbit/editions.ph... granted some predate ISBNs, one is the 3D pop-up version, so not a traditional text, and so forth, but filtering by ISBN will _not_ filter out duplicates. There is also the problem of the same work being published under multiple titles (and also ISBNs) --- Hal Clement's _Small Changes_ was re-published as _Space Lash_ and that short story collection is now collected in: https://www.goodreads.com/book/show/939760.Music_of_Many_Spheres https://www.goodreads.com/book/show/939760.Music_of_Many_Sph... along with others.
- notpushkin 1y agoHmmm, yeah, ISBN isn’t great for this. Is there a good way to deduplicate the books by their contents?
- WillAdams 1y agoLoC or Dewey Decimal with author and title (and edition?) should work. I wish there was some better book cataloging/organizing scheme --- the Online Books Page uses LoC: https://onlinebooks.library.upenn.edu/subjects.html https://onlinebooks.library.upenn.edu/subjects.html and is the most workable of the indices I've used.
- aaron695 1y ago[dead]
- serial_dev 1y agoThe main barriers for me would be: 1. Why? Who would use that? What’s the problem with the other search engines? How will it be paid for? 2. Potential legal issues. The technical barriers are at least challenging and interesting. Providing a service with significant upfront investment needs with no product or service vision that I’ll likely to be sued for a couple of times a year, probably losing with who knows what kind of punishment… I’ll have to pass unfortunately.
- namlem 1y agoIt would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respect international copyright though.
- sam_lowry_ 1y agoLLMs already use it, dude )
- exe34 1y agoI think one use would be to search for information directly from a book, rather than get a garbled/half-hallucinated version of it.
- jdironman 1y agoYou don't need AI for that. I get the optimistic spirit of what you mean though.
- mdp2021 1y agoOptimized information retrieval of complex text is AI.
- echollama 1y agogarbled/half-hallucinated is probably what you would've gotten 8-12mo ago but now adays im sure with good prompting you can pull value from any book.
- tomthe 1y agoI wonder if you could implement it with only static hosting? We would need to split the index into a lot of smaller files that can be practically downloaded by browsers, maybe 20 MB each. The user types in a search query, the browser hashes the query and downloads the corresponding index file which contains only results for that hashed query. Then the browser sifts quickly through that file and gives you the result. Hosting this would be cheap, but the main barriers remain..
- ThatPlayer 1y agoI've done something similar with a static hosted site I'm working on. I opted to not reinvent the wheel, and just use WASM Sqlite in the browser. Sqlite already splits the database into fixed-size pages, so the driver using HTTP Range Requests can download only the required pages. Just have to make good indexes. I can even use Sqlite's full-text search capabilities!
- showerst 1y agoHow would that scale to 10TB+ of plain text though? Presumably the indexes would be many gigabytes, especially with full text search.
- wolfgang42 1y agoThe client only needs to get indexes for the specific search; if the index is just a list of TF-IDF term scores per document (which gets you a very reasonable start on search relevance) some extremely back-of-the-envelope math leads me to guess at an upper bound in the low tens of megabytes per (non-stopword) term, which seems doable for a client to download on demand.
- qcic 1y agoSuper interesting.
- Aachen 1y agoI wonder if you could take this one step further and have opaque queries using homomorphic encryption on the index and then somehow extracting ranges around the document(s) you're interested in Inspired by: "Show HN: Read Wikipedia privately using homomorphic encryption" https://news.ycombinator.com/item?id=31668814 https://news.ycombinator.com/item?id=31668814
- greggsy 1y agoIt's trivial to normalise the various formats, and there were a few libraries and ML models to help parse PDFs. I was tinkering around with something like this for academic papers in Zotero, and the main issue I ran into was words spilling over to the next page, and footnotes. I totally gave up on that endeavour several years ago, but the tooling has probably matured exponentially since then. As an example, all the academic paper hubs have been using this technology for decades. I'd wager that all of the big Gen AI companies have planned to use this exact dataset, and many or them probably have already.
- fake-name 1y ago> It's trivial to normalise the various formats, Ha. Ha. ha ha ha. As someone who as pretty broadly tried to normalize a pile of books and documents I have legitimate access to, no it is not. You can get good results 80% of the time, usable but messy results 18% of the time, and complete garbage the remaining 2%. More effort seems to only result in marginal improvements.
- bawolff 1y ago98% sounds good enough for the usecase suggested here.
- pastage 1y agoWriting good validators for data is hard. You can be 100% sure that there will be bad data in those 98%. From my own experience I thought I had 50% of the books converted correctly and then I found I still had junk data and gave up, it is not an impossible problem I just was not motivated to fix it on my own. Working with your own copies is fine, but when you try to share that you get into legal issues that I just do not feel are that interesting to solve. Edit: my point is that I would like to share my work but that is hard to do in a legal way. That is the main reason I gave up.
- landl0rd 1y ago2% garbage, if some of that garbage falls out the right way, is more than enough to seriously degrade search result quality.
- trollbridge 1y agoDecent storage is $10/TB, so for $10,000 you could just keep the entire 1PB of data. A rather obvious question is if someone has trained an LLM on this archive yet.
- moffkalast 1y agoA rather obvious answer is Meta is currently being sued for training Llama on Anna's archive. You can be practically certain that every notable LLM has been trained on it.
- rthnbgrredf 1y ago> You can be practically certain that every notable LLM has been trained on it. But only Meta was kind of not so smart to publicly admit it.