7 ms·
Running ArchiveTeam's Warrior in Kubernetes
- badlibrarian 2y agoMany of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page into the Wayback Machine every week. Or at least trying to. https://web.archive.org/web/20250122000033/www.google.com https://web.archive.org/web/20250122000033/www.google.com Like so many things about archive.org, when you dig in you start to find wonder and craziness at every turn.
- jfkrrorj 2y ago[flagged]
- badlibrarian 2y agoSuspicion warranted, but citation needed. For now my money is on archives.gov over archive.org. And loc.gov over a website that prioritizes making Pac-Man and Donkey Kong playable in the browser yet leaked the drivers licenses and passports of its patrons, and whose public policy on their wonky javascript UX is "don't read books on a phone."
- jfkrrorj 2y ago[flagged]
- mplewis 2y agoHey man, you only created your account one day ago and you're going to town flooding this site with bad takes. What's the deal?
- Tijdreiziger 2y ago> leaked the drivers licenses and passports of its patrons Source?
- badlibrarian 2y agohttps://www.newsweek.com/catastrophic-internet-archive-hack-hits-31-million-people-1966866 https://www.newsweek.com/catastrophic-internet-archive-hack-... https://www.newsweek.com/internet-archive-hacked-zendesk-1972261 https://www.newsweek.com/internet-archive-hacked-zendesk-197... ----- Subject: Notice of Data Security Incident January 6, 2025 I write on behalf of Internet Archive to inform you about a security incident that involved personal information about you. We regret that this incident occurred and take the security of personal information seriously. On October 20, 2024, we discovered suspicious activity involving our customer service platform. Specifically, between October 17, 2024 and October 20, 2024, an unauthorized actor obtained access to our customer service platform, which contained information about requests from certain Internet Archive users. As soon as we learned of the incident, we took action to contain the incident, including by terminating the unauthorized access, and then launched an investigation to determine the nature and scope of the access. We have determined that the personal information involved in this incident included your name and government ID information such as a driver’s license or passport. As noted above, we took action to contain the incident and investigate it, including by temporarily taking the customer service platform offline. You should always remain vigilant for incidents of fraud and identity theft, including by regularly reviewing and monitoring your accounts. If you discover any suspicious or unusual activity on your accounts or suspect identity theft or fraud, be sure to report it immediately to your financial institutions. If you have been a victim of fraud, you can report it to your local police. Please know that we regret any inconvenience or concern this incident may cause you. Please do not hesitate to contact us at info@archive.org if you have any questions or concerns. Sincerely, Internet Archive
- myself248 2y ago> by proper entities as required by federal law. What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?
- badlibrarian 2y agoSome of the mass deletions are merely a new administration setting up shop. Policies from the previous administration don't belong on the current whitehouse.gov. They wind up here instead https://bidenwhitehouse.archives.gov/ https://bidenwhitehouse.archives.gov/ We pay half a billion in tax dollars for the National Archives, and nearly a billion to the Library of Congress to preserve these records. Others are managed as part of Presidential Libraries. Thousands of employees, dozens of facilities, billions of dollars. Meanwhile archive.org doesn't have air conditioning and preserves physical material within the blast radius of an oil refinery. They let vagrants sleep on their steps yet seem surprised when they set the utility pole outsides on fire. I didn't say it didn't need to be done. I said the whole process needs to be rethought with professional supervision. Setting up more volunteer K8 clusters so that more copies of the Google Home Page can be captured with the wrong user agent isn't going to save democracy.
- toomuchtodo 2y agoArchive.org is outside of the reach of the US government, and is globally distributed. When the US government deletes or darks data (as it has recently done across wide swaths of the federal government website properties), you have no recourse. This means your argument about the resources that go into the US government as a data custodian are meaningless: the outcome is what is material, which is the archival and long term custody & availability of the data sets in scope. Arguably, the Internet Archive has recently proven better at this job than the US government (unsurprising). You're angry at a high value non profit operating on a limited budget. It's weird. I recommend focusing on more important issues than "it is icky around the richmond facility, the power goes out once in a while, and they use ambient air and convection for system cooling which I don't like." If you want to save democracy, the Internet Archive doesn't do that itself. It protects the historical record. If you want to save democracy, that's a different conversation. https://blog.archive.org/2024/05/08/end-of-term-web-archive/ https://blog.archive.org/2024/05/08/end-of-term-web-archive/ https://web.archive.org/collection-search/EndOfTerm2024PreElectionCrawls/ https://web.archive.org/collection-search/EndOfTerm2024PreEl... (no affiliation)
- homebrewer 2y agoHow do I as a non-US citizen get access to information from those "proper entities"? Is it even possible for US citizens? This is often a surprise for some visitors of this fine website, but there's a large world outside the US where "federal law" does not apply.
- badlibrarian 2y agoWe fund the Library of Congress (largest library in the world) and the National Archives (NARA) who make all of this stuff public. Other goverments do similar things. It's all on the web. https://www.archives.gov/presidential-records/research/archived-white-house-websites https://www.archives.gov/presidential-records/research/archi... There are other agencies and data sources to be monitored of course but I'm not seeing a lot of nuance in those efforts yet.
- ch71r22 2y agoFor anyone else interested in running this, it only took a couple seconds to launch their docker-compose.yml https://github.com/ArchiveTeam/warrior-dockerfile/blob/master/docker-compose.yml https://github.com/ArchiveTeam/warrior-dockerfile/blob/maste...
- NortySpock 2y agoI noticed from the docker overlay filesystem that the container was spraying files all over the disk. (Ephemeral, destroyed on container shutdown, sure, but I wanted to reduce write-wear on my ssd...) I tried setting it up with /tmp as a tmpfs (ramdisk) but it then refused to start... Anyone know any broad-spectrum docker incantations to force all overlay writes to RAM, for a container?
- lopkeny12ko 2y ago> the container was spraying files all over the disk Right, that's basically the point...the Warrior downloads files, compresses them, and uploads them for archival. This necessarily requires staging the files somewhere between download and upload. > Anyone know any broad-spectrum docker incantations to force all overlay writes to RAM, for a container? Why would you want this? This sounds like a terrible footgun.
- myself248 2y agoThe Warrior doesn't resume old jobs after a power cycle, so what's the point of committing anything at all to non-volatile storage?
- rustyminnow 2y agoThey say exactly why they want it... "I wanted to reduce write-wear on my ssd"
- j4ah4n 2y agoI think you'll just need to mount it at the right place, with right permissions. Demonstrated here https://stackoverflow.com/questions/39193419/docker-in-memory-file-system https://stackoverflow.com/questions/39193419/docker-in-memor...
- honestSysAdmin 2y ago[dead]
- Havoc 2y agoIsn't there substantial risk involved in having who knows what scraped from your IP?
- tech234a 2y agoYes but many projects are usually restricted to specific websites. A few projects, such as the URLs project, are generally unrestricted.
- WildGreenLeave 2y agoThe first thing I setup when I started to manage my own Kubernetes cluster more then a year ago was this Warrior, I completely forgot about it until this post. Has been active for over a year steadily working the recommended project. Downloaded over 3TB in 6 days (node reboot, so pod was restarted and stats are not persistent). So rough extrapolation is about 180TB. Happy to help the good cause of the ArchiveTeam! Edit: typo