5 ms·
you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the wh
by horseradish7k 1y ago
you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets
- mdaniel 1y agoNo, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org
- viliml 1y agoTangentially related but I've seen twitter links that used to be on the wayback machine disappear from it at some point, presumably due to personal request from the owner.
- corobo 1y agoPretty sure you can nuke all your domains old content by blocking archive.org in robots.txt
- rafram 1y agoI can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.
- mdaniel 1y agoThen why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret? I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings
- rafram 1y agoBecause then you can access the archived destination if you already know the short URL. You just can't get a full list of potentially sensitive short URL/destination pairs.
- yreg 1y agoYeah what they did is probably the best way to handle it.
- mdaniel 1y agoYou are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyrussia21.co.kr which, yes, is for sure very, very secret
- brokensegue 1y agoi asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.
- mdaniel 1y agoThis whole thread is starting to read like some kind of misguided practical joke. I also recognize that it may seem like this is directed toward you, but I'm not shooting the messenger I'm just anchoring my reply under this new information. Sorry about that. But, ok, let's continue in good faith scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive Warrior or its usages scenario 2: they don't want to expose the target of the redirects because it will feed the boundaries of the ravenous AI slurp machines If it's scenario 1, then CSV exists and allows mapping from the 00aa11 codes to the "location:" header, no means and methods necessary If it's scenario 2, then what the hell were they expecting to happen? Embargo the .warc until the AI hype blows over so their great grand children can read about how the Internet was back in the day? I guess the real question is "archive for whom?" because right now unless they have a back-channel way to feed the Wayback Machine's boundary using the .warc files, and thus it secretly populates the Wayback without wholesale feeding the AI boundary, this whole thing is just mysterious