15 ms·
Google Removed 749M Anna's Archive URLs from Its Search Results
- ggm 11mo agoI'm not sure I've ever relied on google to tell me what a site like this had, when the site itself is fully indexed, as this one is. Freetext search over the metastate of title, author, format, date (when available) -seems to work.
- n1xis10t 11mo agoThey don’t have full text search of document contents though do they? I know Google wouldn’t have this for AA pages either, just curious
- ggm 11mo agoGood point. So there is definitely a social utility in search over text which google does have, for the trove it scanned, hands and cats-pawprints and all.
- n1xis10t 11mo agoI’m pretty sure Google indexing pages from Anna’s archive would only get metadata, because AA doesn’t have the full text of the books on those pages. I think to get the full text you have to download the torrents, and I don’t think Google was doing that.
- ggm 11mo agoNo, thats more meta's trick. and they were "only doing it for the articles" not the pictures. I think. I dunno..
- bigiain 11mo agoThey were doing it for the videos too, but only for "personal use"... https://www.wired.com/story/meta-claims-downloaded-porn-at-center-of-ai-lawsuit-was-for-personal-use/ https://www.wired.com/story/meta-claims-downloaded-porn-at-c...
- npteljes 11mo agoWeb searches like Google are great when searching for not exact terms, like synonyms for example. I have never encountered a website that has a search capability like that. Google finds the song "Million voices" by Otto Knows, from the search query "a a a a ah ah ah ah dance song".
- toomuchtodo 11mo agoAre they in ChatGPT and other LLM providers? No need for Google.
- deleted 11mo ago[deleted]
- mmooss 11mo agoThat's a good question: When LLM providers receive DMCA takedowns, how easily can they implement them? Use a post-LLM filter?
- toomuchtodo 11mo agoI was more suggesting that I want my LLM provider to launder the IP so it avoids copyright law. The LLM provider is a fancy search engine where copyright does not apply to the results.
- mmooss 11mo agoDo LLMs filter piracy requests? For example, how will it respond to 'find me a free copy of the Lord of the Rings movies' or more explicitly 'find me a pirated copy ...'?
- shultays 11mo agoProbably yes, I know it at least refuses to 'type down first 5 pages of lotr book' because of copyright reasons. Its filter is getting better (as in worse for the user) everyday
- bean469 11mo ago> how will it respond to 'find me a free copy of the Lord of the Rings movies' or more explicitly 'find me a pirated copy ...' Apparently it depends on the model. Testing on OpenRouter with Search enabled, gpt-5 strictly refuses to provide any links, but Deepseek R1 provides several Archive.org links, one of which is for a torrent file. Thanks Deepseek, I guess I'll be watching The Fellowship of The King for free tonight. ;)
- aunty_helen 11mo agoGoogle does search now? I mean, it's great to see but I'm not sure how this is going to challenge the convenience of my chosen brand of chatbot being able to find the same info without being scammed by 100 seo optimised junk sites.
- JKCalhoun 11mo agoNot sure. I understand they used to do search though. (Love the username, BTW.)
- n1xis10t 11mo agoYeah they’re pretty terrible now. Reminds me, this is an interesting article about search engines getting worse and failing, but the author didn’t get into the spam aspect iirc: https://archive.org/details/search-timeline https://archive.org/details/search-timeline
- zzo38computer 11mo agoIs there a good search engine which does not execute any JavaScripts on files that it scans? (This is not the same as excluding web pages that use JavaScripts (I have seen some search engines that do this); I still want to be able to search for them, but I do not want the search queries (or the summaries of the results) to include anything that is only displayed due to JavaScripts.)
- n1xis10t 11mo agoI haven’t paid attention, so I don’t know. I know that Marginalia lets you downrank or exclude sites that use javascript or something like that, but that isn’t what you’re looking for. You might try Mojeek. I’m sorry I know so little about this.
- n1xis10t 11mo agoI have heard that chatbots aren’t affected by spam as much as Google when you ask them to search, is that true?
- agluszak 11mo agoAnna's archive has already fulfilled G's needs (training Gemini) so now it's time to pretend it never existed ;)
- uvaursi 11mo ago[flagged]
- nine_k 11mo agoDid Anna's Archive also organize much of the world's information and made it universally accessible, for some time?
- GuinansEyebrows 11mo agoThey’re… yes. Yes, that’s exactly what they have done and continue to do. Are you familiar with it?
- someperson 11mo agoFeels weird to say but I have found using Yandex of all places an excellent search engine for content that get taken down by DMCA requests. Eg if you want to watch a movie that's not on Netflix using a web stream the search results are far better. Feels like Google circa 2005.
- negativelambda 11mo agoI just tested, indeed very good results!
- chneu 11mo agoI've been playing around with a variety of search engines such as Kagi, Startpage, Ecosia, DDG. All of them are better than google in finding relevant results. Lol Google is way too "personalized".
- qiqitori 11mo agoYou can turn off personalization. (Operating under the assumption that most people search for facts, I personally don't see why one would ever want personalized results.)
- skulk 11mo ago> I personally don't see why one would ever want personalized results. The same short combination of words can mean very different things to different people. My favorite example of this is "C string" because when I was a kid learning C I was introduced to a whole new class of lingerie because Google didn't really personalize results back then. Now when I search "C string" Google knows exactly what I mean.
- Ariarule 11mo agoI won't bother defending Google-style personalization as it exists for their search results, but since collisions in terminology across fields are common, it's not that hard to see how actual, thoughtful personalization could be useful. Someone searching for "Kafka" is going to want very different results based on whether they're thinking of software or literature. Opinions may also differ over the usefulness of sources, even for people ultimately interested primarily in facts; I find Kagi-style personalization (make your own domain list) very useful, but across Kagi's userbase Reddit is simultaneously one of the most lowered, most raised, and most pinned domains: https://kagi.com/stats?stat=leaderboard https://kagi.com/stats?stat=leaderboard
- drnick1 11mo agoGo thing that Google hasn't been a part of my life for a while now. I use DuckDuck for search.
- NooneAtAll3 11mo agoI've seen DDG censor stuff that was still on google
- aucisson_masque 11mo agoDuckduckgo is bing, bing is Microsoft. I don't see how Microsoft is better than google at censorship.
- culi 11mo agoddg actually has its own crawler and does a tiny amount of its own indexing. It used to do more but resorted to just mostly using Bing and Yandex indexes
- storus 11mo agoGoogle's march to irrelevance continues with full steam.
- DaSHacka 11mo agoThey got a long way ahead of them then, considering they're still something like 97% of all search queries.
- esafak 11mo agoActually ~90%, but that does not include AI search (chatgpt et al). https://www.klatch.co.uk/search-engine-market-share https://www.klatch.co.uk/search-engine-market-share
- chris_wot 11mo agoGoogle search keeps getting less useful every day.
- ilt 11mo agoAnd still it’s the top result in Google if one searches for Anna’s archive. How is it that that search result hasn’t been removed?
- incompatible 11mo agoPresumably, the home page doesn't contain any copyright violations. This is only DMCA stuff targetting individual links.
- tonyhart7 11mo agowell if publisher DMCA request to google then I don't know why people get mad about its still piracy at the end of the day and publisher have right to license etc, people mad about this maybe dont have to deal this as a business
- pessimizer 11mo agoI was surprised that those pages showed up in book title searches at all. Makes sense to get rid of them, you don't want a search for a book to be topped by a link to pirate the book. The top-level domains still come up, and people who know they want to pirate a book can still find the site.
- almaight 11mo agohttps://www.google.com/search?q=Anna%27s+Archive https://www.google.com/search?q=Anna%27s+Archive
- deleted 11mo ago[deleted]
- musicale 11mo agoGoogle has already removed URLs from the first page of "search" results.
- nullbyte808 11mo agohttps://annas-archive.org https://annas-archive.org
- nullbyte808 11mo agoMan I need to get around to downloading the z-archive torrents before annas archive is taken down. If I eliminate large PDFs and non english books I think I can fit it on two 32 TB drives with BTRFS z-std compression max setting. https://annas-archive.org/torrents https://annas-archive.org/torrents
- cookiengineer 11mo agoLet me know of those efforts, I wanna have an English/German/French backup of the archive, too. But as you said HDDs and filesystems are the problem, really. Maybe I'll have to build a torrent splitter or something, because the UIs of all torrent clients are just not built for that.
- h4ck_th3_pl4n3t 11mo agoSneed
- mmooss 11mo ago> eliminate large PDFs How large? Isn't that going to result in an arbitrary filter of books? In other domains, large PDFs are due to PDF production errors, such as using color or needlessly high resolution, and not so much due to the volume of content - at least for text.
- Llamamoe 11mo agoDepending on how important it is for you to maintain original quality, I have in the past had good luck with a combination of prerendering complex content, reducing the DPI and colour depth of images, and recombining them back into PDFs, depending on the file. You could probably easily automate identifying different editions of the same content, and e.g. only keep an epub with small images, rather than the other 6 and 3 more PDFs as well.
- brador 11mo agoInvert the list, start with the smallest, continue until full.
- 0xedd 11mo ago[dead]
- jimjimwii 11mo agoI am not exaggerating when i say i completely stopped using google for searches that google might take offence to. Serial numbers, business phone numbers, and of course books and papers all ho through real search engines. Currently, those are yandex as my main goto with brave as a backup. I couldn't care less what google does because i don't use it.
- Razengan 11mo agoWait so did Gemini train on Wikipedia etc.? Isn't it a conflict of interest or something if their AI results prevent people from clicking on the websites Google's AI trained on?
- renegat0x0 11mo agoSearching the web has changed: - There are more walled gardens, so engines legally cannot enter some spaces - There are more legal problems with data, so more things are not accessible - to find stuff you have to check google, but also yandex, or kagi, or chatgpt - I also check my own index for stuff https://github.com/rumca-js/Internet-Places-Database https://github.com/rumca-js/Internet-Places-Database
- aswegs8 11mo agoOn a related note, I think Anna's archive might be the last remaining bastion for books after library genesis got shut down recently. Is anyone aware of other alternatives?
- chakintosh 11mo agoWeLib.org for books AudiobookBay for audiobooks
- culi 11mo agoIs WeLib an Anna's Archive mirror? Seems very similar. I still mainly use LibGen for books. Got me through college and probably saved me well over $2k on textbooks throughout my courses
- rendx 11mo agoLinked from Anna's Archive: https://open-slum.org/ https://open-slum.org/
- culi 11mo agoAt least for academic papers, the network is still around but has moved to a more decentralized solution. Nowadays, the bleeding edge is a network of [mostly Telegram] bots that you give a doi to and they return your desired paper. It's called Nexus (or LibrarySTC?) https://libstc.nexus/ https://libstc.nexus/ It's very fast and efficient. I've never seen a bot get taken down either.
- dev1ycan 11mo agoOh wow just what I said would happen, happened... first libgen and z-lib after META trained its model with 70tb of torrented content and now Anna's library. Meanwhile REAL human students and researchers lose access to acadeemic work
- rodolphoarruda 11mo agoA question to the community: would it be a (legal) problem if I decided to download digital copies of the physical books I already have in my bookshelf? I was thinking on using Anna's Archive for that. Hobby project.
- syntaxers 11mo ago17 USC 106 gives copyright holders exclusive rights to reproduce and distribute copies; no exemption exists for downloading digital copies because you own the physical book, and fair use (17 USC 107) is unlikely to apply when commercial alternatives exist and you’re copying entire works from unauthorized distributors.
- rodolphoarruda 11mo ago> you’re copying entire works from unauthorized distributors Yep, this sounds like an issue. So the idea from MP3 early days of "let me download these files as a backup before I lend my CD collection to my cousin" is not a real option.
- probably_wrong 11mo agoAs far as my extremely poor understanding of the law goes: this depends on where you live but generally you are not allowed to download a digital copy of a physical book you own, but you are allowed to create your own [1]. It may also be worth noting that most jurisdictions are only interested in distribution, not downloading, so the chances of prosecution are slim. A small company you may have heard of called Meta is currently using a similar argument in US court [2]. [1] https://ebooks.stackexchange.com/questions/1111/i-have-a-printed-version-of-a-book-does-it-allow-me-to-possess-an-electronic-co https://ebooks.stackexchange.com/questions/1111/i-have-a-pri... [2] https://news.ycombinator.com/item?id=43125840 https://news.ycombinator.com/item?id=43125840
- fedeb95 11mo agono problem, AA has a very good search bar.
- extraduder_ire 11mo agoDoes google still link to lumendatabase.org (formerly chillingeffects) when results have been taken down due to a legal request?
- submeta 11mo agoGoogle also has deleted hundreds of videos on Youtube documenting Israel's crimes in Gaza. So did X: Remove thousands of videos and accounts documenting Israel's war crimes in Gaza. These companies are evil. Will always side with the strong and powerful.
- MarsIronPI 11mo agoAt this point someone could make a piracy search engine that crawls all these reported URLs.
- culi 11mo agoYandex basically does this already tbh
- pacman1337 11mo agoIf you don't have access to massive amounts of digitized books you are at a significant competitive disadvantage i.e, AI + RAG is a game changer for consuming technical content. That last piece of the puzzle I am missing for my setup is being able to digitize the books as markdown + latex for mathematics equations, right now it is just expensive.