5 ms·
The main barriers for me would be: 1. Why? Who would use that? What’s the problem with the other search engines? How will it be paid for? 2. Potential legal i
by serial_dev 1y ago
The main barriers for me would be:
1. Why? Who would use that? What’s the problem with the other search engines? How will it be paid for?
2. Potential legal issues.
The technical barriers are at least challenging and interesting.
Providing a service with significant upfront investment needs with no product or service vision that I’ll likely to be sued for a couple of times a year, probably losing with who knows what kind of punishment… I’ll have to pass unfortunately.
- namlem 1y agoIt would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respect international copyright though.
- sam_lowry_ 1y agoLLMs already use it, dude )
- exe34 1y agoI think one use would be to search for information directly from a book, rather than get a garbled/half-hallucinated version of it.
- jdironman 1y agoYou don't need AI for that. I get the optimistic spirit of what you mean though.
- mdp2021 1y agoOptimized information retrieval of complex text is AI.
- echollama 1y agogarbled/half-hallucinated is probably what you would've gotten 8-12mo ago but now adays im sure with good prompting you can pull value from any book.
- jxjnskkzxxhx 1y agoDo you have a reason to believe this ain't already being done? I would assume that the big guys like openai are already training on basically all text in existence.
- IlikeKitties 1y agoIn fact, facebook torrented annas archive and got busted for it, because of course they did: https://torrentfreak.com/meta-torrented-over-81-tb-of-data-through-annas-archive-despite-few-seeders-250206/ https://torrentfreak.com/meta-torrented-over-81-tb-of-data-t...
- HDThoreaun 1y agoEvery LLM maker probably did the same. Facebook just has disgruntled employees who leaked it
- gpm 1y agoGoogle goes around legally scanning every book they can get their hands on with books.google.com. Legally scanning every paper they can get their hands on with scholar.google.com. I doubt they'd resort to piracy for what is basically the same information as what they've already legally acquired...
- lcnPylGDnU4H9OF 1y agoThat is a good reason to think they did not but it doesn't necessarily override reasons for them to do so. Perhaps it's dubious that the subset of data they could not legally get their hands on is an advantage for training but I really don't know, and maybe nobody does. Given that, Google's execs may have been in favor of similar operations as Facebook's and their lawyers may have been willing to approve them with similar justifications.
- sneak 1y agoDownloading a torrent isn't piracy if you are a license holder for the information that you are downloading.
- executesorder66 1y ago> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.
- TeMPOraL 1y agoThere's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.
- corgi912 1y agoThere's this famous phrase in Russian that was born out of a short interview with a woman, a strong Putin supporter, that's often been used as a sarcastic remark for pointing out someone's double standards and/or hypocrisy. It can be roughly translated to "you don't understand, it's a completely different situation". That's what's constantly on my mind when I'm reading discussions like this one. Everybody and their dog torrenting petabytes of data and getting away with it (Meta is the only one that got caught and they've still gotten away with doing it)? The very same data poor American students were forced to commit suicide over? The same data that average American housewives were sued over for millions of dollars of "damages"? The same data that often gets random German plumbers or steelworkers to pay thousands of euros of "fines" to the copyright mafia so they won't get sued and have their lives ruined? Yet when giant corporations are doing the exact same thing on a massive scale, it's fine? It's not even the same thing, an American student torrenting books isn't making any money off it, while Meta very much is. Of course it's not the same, a simple-minded and poorly educated person like me isn't capable of understanding the difference. You keep believing in your moral superiority, the rest of the world has finally woken up.
- Exoristos 1y agoThere are those who are in charge and those who aren't.
- andrepd 1y ago> Would probably have to be done in Russia or some other country that doesn't respect international copyright though. Incredible, several years of major American AI companies showing that flaunting copyright only matters if it's college kids torrenting shows or enthusiasts archiving bootlegs on whatcd, but if it's big corpos doing it it's necessary for innovation. Yet some people still believe "it would have to be done in evil Russia".
- deleted 1y ago[deleted]
- DataDaoDe 1y agoOP does have an exaggerated statement - its not like there aren't laws in Russia or something and I largely agree with your sentiment. I think there are levels to this though and its pretty clear that Russia is much riskier than the USA when it comes to IP - just look up anything to do with insuring IP risk in Russia (here's one such example: https://baa.no/en/articles/i-have-ip-in-russia-is-my-ip-at-risk https://baa.no/en/articles/i-have-ip-in-russia-is-my-ip-at-r...) Also according to the office of US trade representative, Russia is on the priority watch list of countries that do not respect IP [1] and post 2022, largely due to the war, Russia implemented measures negatively effecting IP rights. [2,3] If you think it isn't the case and Russia is just as risky as the US when it comes to copyright and IP, I would be interested to know why. 1. https://ustr.gov/about/policy-offices/press-office/press-releases/2025/april/ustr-releases-2025-special-301-report-intellectual-property-protection-and-enforcement#:~:text=Eight%20countries%20are%20on%20the,engagement%20during%20the%20coming%20year. https://ustr.gov/about/policy-offices/press-office/press-rel... 2. https://www.papula-nevinpat.com/executive-summary-the-ip-situation-in-russia-and-ukraine/ https://www.papula-nevinpat.com/executive-summary-the-ip-sit... 3. https://www.taftlaw.com/news-events/law-bulletins/russia-issues-decree-affecting-ip-rights-for-unfriendly-countries/?utm_source=chatgpt.com https://www.taftlaw.com/news-events/law-bulletins/russia-iss...
- mdp2021 1y ago> evil In this case and context, a label like "evil" is a twisted interpretation.
- carlosjobim 1y ago> 1. Why? Who would use that? Rather who would use a traditional search engine instead of a book search engine, when the quality of the results from the latter will be much superior? People who need or want the highest quality information available will pay for it. I'd easily pay for it.
- bbor 1y ago1. It'd be for the scientific community (broadly-construed). Converting media that is currently completely un-indexed into plaintext and offering a suite of search features for finding content within it would be a game-changer, IMO! If you've ever done a lit review for any field other than ML, I'm guessing you know how reliant many fields are on relatively-old books and articles (read: PDFs at best, paper-only at worst) that you can basically only encounter via a) citation chains, b) following an author, or c) encyclopedias/textbooks. 2. I really don't see how this could ever lead to any kind of legal issue. You're not hosting any of the content itself, just offering a search feature for it. GoodReads doesn't need legal permission to index popular books, for example. In general I get the sense that your comment is written from the perspective of an entrepreneur/startup mindset. I'm sure that's brought you meaning and maybe even some wealth, but it's not a universal one! Some of us are more interested in making something to advance humanity than something likely to make a profit, even if we might look silly in the process.
- Aachen 1y ago> I really don't see how this could ever lead to any kind of legal issue. You're not hosting any of the content itself, just offering a search feature for it. You don't need to host copyrighted material. It's all about intent. The Pirate Bay is (imo correctly, even if I disagree with other aspects about copyright law and its enforcement) seen as a place where people go to find ways to not pay authors for their content. They never hosted a copyrighted byte but they're banned in some form (DNS, IP, domain seizures) in many countries. Proxies of TPB also, so being like an ISP for such a site is already enough, whereas nobody is ordering blocks of Comcast's IP addresses for providing access to websites with copyrighted material because they didn't have a somewhat-provable intent to provide copyright infringement When I read the OP, I imagine this would link from the search results directly to Anna's archive and sci-hub, but I think you'd have to spin it as a general purpose search page and ideally not even mention AA was one of the sources, much less have links (Don't get me wrong: everyone wants this except the lobby of journals that presently own the rights) It would be a real shame if an anonymous third party that's definitely not the website operator made a Firefox add-on that illegitimately inserts these links to search results page though
- 1vuio0pswjnm7 1y agoBut he did not mention anything about creating a "service" It could be his own copy for personal use What if computers continue to become faster and storage continues to become cheaper; what if "large" amounts data continue to become more manageable The data might seem large today, but it might not seem large or unmanageable in the future