7 ms·
One of the things I dislike the most about Internet Archive is their relatively open attitude of flaunting privacy on purpose: http://blog.archive.org/2017/04/1
by _ktx2 4y ago
One of the things I dislike the most about Internet Archive is their relatively open attitude of flaunting privacy on purpose: http://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ http://blog.archive.org/2017/04/17/robots-txt-meant-for-sear...
Do you know what system they replaced robots.txt with? Email, one that is filed as a DMCA request: https://medium.com/wednesday-genius/how-to-remove-your-website-from-the-internet-archive-2020-c4d89c147546 https://medium.com/wednesday-genius/how-to-remove-your-websi... https://jonathanwthomas.net/how-to-get-your-website-out-of-the-internet-archive-wayback-machine/ https://jonathanwthomas.net/how-to-get-your-website-out-of-t...
Sometimes, it's probably good to not push the envelope without trying to establish consensus in good faith first.
- rasz 4y agoIsnt that a response to companies buying old unused domains, slapping robots on it and thus killing whole archive of this domain going back 20 years?
- _ktx2 4y agoCould be! However, making a direct attack on individual privacy should never have been an option. To make matters worse, the logic of, "We did this to government and military websites, so now we're going to roll it out everywhere" was quite broken for the time and remains so. There's examples of how this works in a healthy way. Martin Manley is one scenario that comes to mind, where he overtly opted-in to having an archive stored about him upon his death: https://martin-manley.eprci.com/ https://martin-manley.eprci.com/
- daniel_reetz 4y agoWith respect, I fail to see how a public website is a privacy matter.
- stubish 4y agoInformation on a public website is public until it is taken down or the information changed. The Internet Archive removes an individuals control over when the information remains public. This is privacy. We might be caught naked, and we can't unsee what has been seen, but it is a basic human instinct to draw the curtains and contain further damage. Perfectly innocent individuals suffer because the IA rules are designed around edge cases where public figures try to hide misdeeds.
- bakugo 4y ago> The Internet Archive removes an individuals control over when the information remains public. And that's a good thing in the vast majority of cases. Unless we're talking about sensitive information that was published without the consent of the person in question, all public information should remain public forever.
- stubish 4y agoIn my experience, it is the vast minority of cases. Most of the content of the IA is not in the public interest, now or in the future. It is crap. It is noise. It is the contents of the Internet at a point in time. Actual information is the wheat in the chaff, and why you need search engines to find it. We know this, because of the Usenet archives that are intermittently available. Almost completely useless apart from people having a giggle at how the Internet used to be, a quick browse and search for naughty words. And a few gems in the mountain of noise, in such dire need of curation people hardly know it exists and barely justifiable enough for libraries to keep it alive.
- gojomo 4y agoAgreed, bulk collection gets dominated by crap, which individually has little value. But there's some absolutely essential priceless diamonds hidden in the crap. And they can't be found/known at the time of collection: only with the future development of other events & knowledge do they become retroactively evident. So you've got to collect & preserve as much as you practically can, or else great things are lost forever. Further, even the mounds/magnitudes of crap can turn out to be important for understanding the past. Ads that annoyed readers at the time help communicate how people, & businesses, & technology were really operating – not just the self-serving stories people craft later. The most-fumbling and awkward early uses of a new medium – hypertext, or RealAudio, or Shockwave Flash, or whatever – reveal enduring lessons about the evolution of technology & culture, including roads-not-taken that could still hold promise. This shouldn't surprise us. Much of what we know of past civilizations comes from archeologists studying trash dumps that, via dumb luck, were well-preserved. So if you tell me, "the Wayback Machine is a giant unedited trash heap of the internet", my response is: "Yes! That's the point! You get it!"
- gojomo 4y agoNeiher 'flaunting privacy' nor 'direct attack on individual privacy' are fair descriptions of any of the Archive's web collection policies. People who freely publish information, to the worldwide public, on the 'World Wide Web' should reasonably expect all sorts of entities to collect, save, analyze, & repurpose that info, unless they take specific steps to discourage such access & use. The Archive's crawlers identify themselves, and collect things that are publicly linked, or specifically nominated-for-collection by library patrons or partners. Except in some focused specialized collection projects, they don't "log in" as any user, only visiting & collecting what's published freely to any anonymous person/organization/process. For material needing more privacy, websites always have the option to block any and all unwanted visitors/crawlers with a wide variety of standard techniques, like requiring logins or simple challenges that automated crawlers won't pass. And, as your linked articles report, the process for a later exclusion by request is pretty quick and simple. (The 2nd post concludes: "So, hats off to the Internet Archive for making the process smooth and relatively painless.") And, such exclusion does not require any sort of "DMCA request".
- stubish 4y agoThis is victim blaming. In my jurisdiction, you retain copyright under any information you publish, even to the worldwide public. This means I can reasonably expect entities to collect, save, analyze and repurpose that info within reason, and without specific steps to discourage access & use. This is why there are laws such as 'fair use' and 'satire', because we wanted to extend what is considered reasonable use of public works. But redistributing copyrighted works without permission? Legally actionable, if you have the money and lawyers and access to the necessary courts. If this was software, such as free software license violations, people in this forum would be calling for the lawyers to nuke them from orbit. Thankfully DMCA should make the removal process easier now, especially in situations where control over the domain has been lost or being hosted by a third party. Although last I saw there were still artificial barriers, such as needing to list every single individual page needing to be taken down. But this is after the fact, after you discovered your reasonable expectations and privacy have been violated. And then you have to track down the other copies that IA illegally distributed your now-private and copyrighted information to, such as a few libraries around the world with similar projects.
- tornato7 4y agoIf you’re relying on an honor system .txt file to preserve your privacy I think that says enough already. It’s not like they’re infiltrating password-protected links or private iCloud accounts.
- _ktx2 4y agoYou are right, regulation is what solves this for good
- gojomo 4y agoIn truth, regulation can only reliably protect your privacy from well-behaved actors whose actions/violations are observable. If you've taken no self-help measures to limit access, then bad actors, unobservable to you and regulators, will still be doing whatever they would like to do and can get away with. But you may be lulled into a false sense of security by the false promise of a 'solution' via regulations.
- _ktx2 4y agoAs I've now learned, you used to work for the Internet Archive. You should probably start your statements with that. > If you've taken no self-help measures to limit access... robots.txt was a nice self-help measure. > ... then bad actors, unobservable to you and regulators, will still be doing whatever they would like to do and can get away with. Regulators still have to follow regulations. You are right that I can't stop someone from creating offline archives - but they're not really who I am worried about. Nor am I worried about the small servers that keep copies of documents during transmission, unless of course they're doing so for criminal reasons.
- gojomo 4y agoShould you start all of your statements with a list of every project you've ever worked on? Show me an example of how it's done before you make such an exceptional request of me. For any who are more curious about a commenters' background than their current words, my profile already links to copious resources on my work history, & writings, beyond what's typical of contributors here.