9 ms·
ArchiveTeam has finished archiving all goo.gl short links
- brador 1y agoGamefaqs remains unarchived.
- mdaniel 1y agoBe the change you want to see in the world. Contributing .warc files to Archive.org isn't a gated club. My understanding of calling down the Warrior team is when something is time sensitive and needs to pseudo-ddos the site to get the bytes right now. Unless you know something about the demise of Gamefaqs, you have the rest of your life to archive a page at a time
- do_not_redeem 1y agoDoes "all" mean all the URLs publicly known, or did they exhaustively iterate the entire URL namespace?
- jedberg 1y agoThey iterated the entire URL namespace by having volunteers run a client so they didn't get IP banned.
- barbazoo 1y agoBeautiful. I wish I had seen this and could have helped.
- brokensegue 1y agothey are still archiving other url shorteners https://tracker.archiveteam.org:1338/ https://tracker.archiveteam.org:1338/ you can participate in that
- Imustaskforhelp 1y agoare we sure that the whole entire URL namespace has been mapped? How would that even function, I mean, did they loop through every single permutation and see the result, or what exactly/ how would that work?
- toomuchtodo 1y agoThe pipeline code is available for review of the mechanics of http requests made if you follow the ArchiveTeam wiki links.
- jedberg 1y ago> did they loop through every single permutation and see the result, or what exactly/ how would that work? In short, yes. Since no one can make new links, it's a pre-defined space to search. They just requested every possible key, and recorded the answer, and then uploaded it to a shared database.
- deleted 1y ago[deleted]
- ccgreg 1y agoThe goo.gl URLs that are publicly known are already in the Internet Archive and Common Crawl crawls.
- zdimension 1y agoTitle is imprecise, it's Archiveteam.org, not Archive.org. The Internet Archive is providing free hosting, but the archival work was done by Archiveteam members.
- im3w1l 1y agoWhat exactly is archiveteam's contribution? I don't fully understand. Edit: Like they kinda seem like an unnecessary middle-man between the archive and archivee, but maybe I'm missing something.
- 1gn15 1y agoArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)
- notpushkin 1y agoSidenote, but you can also run a Warrior in Docker, which is sometimes easier to set up (e.g. if you already have a server with other apps in containers).
- kalleboo 1y agoYep, I have my archiveteam warrior running in the built-in Docker GUI on my Synology NAS. Just a few clicks to set up and it just runs there silently in the background, helping out with whatever tasks it needs to.
- gunalx 1y agoRan archive warrior a while back but hadde to shut it down AS i sterted seeing the VM was compromised trying to spam ssh and other login attemps in my local network.
- 1y ago
- Ayesh 1y agoRecent update from Google: https://blog.google/technology/developers/googl-link-shortening-update/ https://blog.google/technology/developers/googl-link-shorten...
- OJFord 1y agoThis leaves me wondering what the point is? What could it possibly cost to keep redirecting existing shortlinks that they consider unused/low activity already anyway? (In addition to the higher activity ones parent link says they'll now continue to redirect.)
- toomuchtodo 1y agoTo save face.
- RicoElectrico 1y agoIn another submission someone speculated the reason might be the unending churn of the Google tech stack that just makes low-maintenance stuff impossible.
- immibis 1y agoMy guess is that plus not having a single person left to maintain it due to the similarly unending people churn.
- manquer 1y agoFor a company also running a hosting service like GCP? nothing. They already have plenty of unused compute /older hardware / CDN POPs, performant distributed data store and everything else possibly needed . It would be cheaper than the free credits they giveaway just one startup to be on GCP. I don’t think infra costs are a factor in a decision like this .
- shaky-carrousel 1y agoYeah, I'll take that "update" like the extremely unreliable info from an extremely unreliable company that it is.
- dang 1y agoRelated. Others? Enlisting in the Fight Against Link Rot - https://news.ycombinator.com/item?id=44877021 https://news.ycombinator.com/item?id=44877021 - Aug 2025 (107 comments) Google shifts goo.gl policy: Inactive links deactivated, active links preserved - https://news.ycombinator.com/item?id=44759918 https://news.ycombinator.com/item?id=44759918 - Aug 2025 (190 comments) Google's shortened goo.gl links will stop working next month - https://news.ycombinator.com/item?id=44683481 https://news.ycombinator.com/item?id=44683481 - July 2025 (222 comments) Google URL Shortener links will no longer be available - https://news.ycombinator.com/item?id=40998549 https://news.ycombinator.com/item?id=40998549 - July 2024 (49 comments) Ask HN: Google is sunsetting goo.gl on 3/30. What will be your URL shortener? - https://news.ycombinator.com/item?id=19385433 https://news.ycombinator.com/item?id=19385433 - March 2019 (14 comments) Tell HN: Goo.gl (Google link Shortener) is shutting down - https://news.ycombinator.com/item?id=16902752 https://news.ycombinator.com/item?id=16902752 - April 2018 (45 comments) Google is shutting down its goo.gl URL shortening service - https://news.ycombinator.com/item?id=16722817 https://news.ycombinator.com/item?id=16722817 - March 2018 (56 comments) Transitioning Google URL Shortener to Firebase Dynamic Links - https://news.ycombinator.com/item?id=16719272 https://news.ycombinator.com/item?id=16719272 - March 2018 (53 comments)
- makeworld 1y agoGlad I contributed to this in some small way.
- Klathmon 1y agoSame, it's nice to see my username on the leaderboards. Even though all I did was setup the docker container one day and forget about it
- yreg 1y agoI wonder how many of them lead to private YouTube videos, Google documents, etc.
- mdaniel 1y agoI was going to be cheeky and say "well, now you can download them and search" but it seems it's "Access-restricted-item: true" for some reason, above and beyond being 10G a pop <https://archive.org/details/archiveteam_googl_20250228144231_a8e742fa https://archive.org/details/archiveteam_googl_20250228144231...>
- horseradish7k 1y agoyou'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets
- mdaniel 1y agoNo, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org
- viliml 1y agoTangentially related but I've seen twitter links that used to be on the wayback machine disappear from it at some point, presumably due to personal request from the owner.
- corobo 1y agoPretty sure you can nuke all your domains old content by blocking archive.org in robots.txt
- rafram 1y agoI can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.
- dkh 1y agoExcellent! ArchiveTeam have always been impressive this way. Some years ago, I was working at a video platform that had just announced it would be shutting down fairly soon. I forget how, but one way or another I got connected with someone at ArchiveTeam who expressed their interest in archiving it all before it was too late. Believing this to be a good idea, I gave them a couple of tips about where some of our device-sniffing server endpoints were likely to give them a little trouble, and temporarily "donated" a couple EC2 instances to them to put towards their archiving tasks. Since the servers were mine, I could see what was happening, and I was very impressed. Within I want to say two minutes, the instances had been fully provisioned and were actively archiving videos as fast as was possible, fully saturating the connection, with each instance knowing to only grab videos the other instances had not already gotten. Basically they have always struck me as not only having a solid mission, but also being ultra-efficient in how they carry it out.
- Aardwolf 1y agoI don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?
- immibis 1y agoThey might be storing in WARC format, which records all the request and response headers and maybe even TLS certificates and things.
- jdiff 1y agoI did some ridiculous napkin math. A random URL I pulled from a Google search was 705 bytes. A googl link is 22 bytes but if you only store the ID, it'd be 6 bytes. Some URLs are going to be shorter, some longer, but just ballparking it all, that lands us in the neighborhood of hundreds of billions of URLs, up to trillions of URLs.
- rafram 1y ago> A random URL I pulled from a Google search was 705 bytes. 705 bytes is an extremely long URL. Even if we assume that URLs that get shortened tend to be longer than URLs overall, that’s still an unrealistic average.
- jdiff 1y agoIt is long, it represents the lower hundreds of billions bound in my awful napkin math.
- digitaldragon 1y agoThe data is saved as a WARC file, which contains the entire HTTP request and response (compressed, of course). So it's much bigger than just a short -> long URL mapping.
- deleted 1y ago[deleted]
- SilverElfin 1y agoIs there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.
- 9dev 1y agoAsk OpenAI maybe?
- DaSHacka 1y ago> reddit There used to be one such project (Pushshift), before the Reddit API change. You can download all the data and see all the info on the-eye, another datahoarder/preservationist group: https://the-eye.eu/redarcs/ https://the-eye.eu/redarcs/ > twitter Not that I know of, and you haven't even been able to archive tweets on the Wayback machine for YEARS.
- stuffoverflow 1y agoAcademictorrents has monthly dumps of all reddit submissions and comments even after the API restrictions.
- pabs3 1y agohttps://academictorrents.com/browse.php?search=stuck_in_the_matrix%2C+Watchful1%2C+RaiderBDev https://academictorrents.com/browse.php?search=stuck_in_the_...
- SilverElfin 1y agoInteresting. You don’t have to be an academic to access these I guess?
- mkl 1y agoThey have magnet links and torrent files right there on the pages, so no.
- Seattle3503 1y ago
- iJohnDoe 1y agoWhy? Did they ask anyone if it was okay? Anything sensitive at those links? Anything at those links people didn't want or need anymore? Maybe people thought those links were dead? Did Google provide a way to cancel those links first? It's like when the GPT links were archived and publicly available that contained sensitive information.
- diath 1y agoIf you want something to remain private, don't post it on the public internet.
- wiredpancake 1y agoSometimes to preserve history, you just have to go do what you gotta do. After all, these are just short links. They link to other things on the Internet. Which is inherently public anyways. You cannot expect privacy via a simple URL. These short URLs are short, hence programmatically scraping all the URLs. The GPT Links situation is nothing like this imo. Both however do come down to the stupid human aspect.
- anticrymactic 1y agoIt's a link, what privacy can one expect? Especially with short links there's always the possibility of entering ~6 characters and getting a hit. So I believe expecting any secrecy from urls is silly. That's like posting your passwords on Twitter because "Why would anyone find my account"
- NylaTheWolf 1y agoHell yeah!!! Fantastic work, everyone!
- m3kw9 1y agoOk how do I access them, or is that not the point?
- raldi 1y agoGoogle said they would keep hosting any recently-clicked link; does this mean that all the links are now recently-clicked?
- JimDabell 1y ago“Recently clicked” wasn’t the criterium, it was “showed activity in late 2024”. So nothing that anybody has done this year – including this archiving – will affect which links Google keep alive.
- raybb 1y agoHappy go have contributed a hundred thousand links by running their docker container!
- edg5000 1y agoCan we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot measures over time on websites, in turn making it easier to crawl. Does something like this exist? It would be so awesome. It would also allow people to run a search engine at home.
- bayindirh 1y agoWhy would I spend time and resources to feed a machine which wastes more resources to hallucinate fiction from data it ingested? For digital preservation? We may discuss. For an LLM? Haha, no. No, thank you.
- lyu07282 1y ago> This would mean there is an "official" source of all web data. LLM people can use snapshots of this that already exists, its called CommonCrawl: https://commoncrawl.org/ https://commoncrawl.org/
- patrickhogan1 1y agoCommon Crawl, while a massive dataset of the web does not represent the entirety of the web. It’s smaller than Google’s index and Google does not represent the entirety of the web either. For LLM training purposes this may or may not matter, since it does have a large amount of the web. It’s hard to prove scientifically whether the additional data would train a better model, because no one (afaik) not Google not common crawl not Facebook not Internet Archive have a copy that holds the entirety of the currently accessible web (let alone dead links). I’m often surprised using GoogleFu at how many pages I know exist even with famous authors that just don’t appear in googles index, common crawl or IA.
- schoen 1y agoIs there any way to find patterns in what doesn't make it into Common Crawl, and perhaps help them become more comprehensive? Hopefully it's not people intentionally allowing the Google crawler and intentionally excluding Common Crawl with robots.txt?
- udir_net 1y ago[dead]