3 ms·
I'll take a contrarian view and probably get downvoted for it, but many people are adamant about the benefits of indiscriminate archiving, and I just don't see
by zerobees 3mo ago
I'll take a contrarian view and probably get downvoted for it, but many people are adamant about the benefits of indiscriminate archiving, and I just don't see it. Do we have a moral right to keep a copy of everything that's ever been written on the internet, basically just for the sake of it?
Sure, there's a variety of official and quasi-officials resources that should be treated as public record and preserved. And arguably, there are things that rise to the level of a cultural phenomenon and where the benefit of keeping receipts outweighs the jerk factor of never asking for permission and not respecting the wishes of private individuals.
But if it's some family blog from 30 years ago that's been deliberately taken down and lives on archive.org unbeknownst to the original owner? Do we have a right to that? To what end, other than "well, future historians may need it"? A historian won't look at it. A person trying to doxx you or shame you will.
- CalRobert 3mo agoThat's a good point, and AI makes this worse unfortunately. I realized a while ago I spend more time taking pictures than I even do looking at them, and quit worrying so much about saving everything. As they say, all these moments will be lost in time, like tears in rain. On the other hand, it's a memory of an internet that no longer exists, and that I think a lot of people (myself included) miss dearly. I do suspect a lot of things weren't deliberately taken down so much as just not maintained, though.
- badlibrarian 3mo agoArchives have not taken a consistent stance on this. Preservation, removal, restricted access, de-indexing, and “right to be forgotten” sit on a spectrum. Good archival practice has to include judgment, context, and humility. Sometimes that means preserving. Sometimes it means limiting access. And sometimes it may mean "honoring" a removal request or court order, even if you're just setting a flag.
- jjav 3mo agoIn the pre-digital era, the norm was that content was preserved. Simply because it was printed in newspapers, magazines, books, alsbums, etc.. in thousands (or hundreds of thousands) of copies, so merely by inertia, many copies survived squirreled away in libraries, attics and other storage. Also, once printed and distributed, it couldn't be taken back and altered since countless original copies existed. Sure, a lot of it wasn't that important, but only in hindsight of history does it become apparent what was important. So it was possible to go back and research those old unaltered originals. I fear for history in the digital era. Everything is fleeting, everything is erased and everything can be retroactively altered when the powers-that-be don't like what was said. Thinking centuries ahead, reliable historical records basically stop around 2000-2020 or so.
- ninjalanternshk 3mo ago> In the pre-digital era, the norm was that content was preserved. So pre-digital, no books, publications, photographic prints, scrolls, tablets, or clay etchings we lost to time? > Thinking centuries ahead, reliable historical records basically stop around 2000-2020 or so. That’s basically backwards. We’re making a stink because people are identifying one needle in a giant pile of needles that they can’t find anymore. And people point to thousand-year old physical writings that persist, because they persist, and don’t know what was lost because it was lost. Digital records are far far more long-lived — in the statistical average — than physical.
- amazing_stories 3mo agoI'm not sure I've read anything more wrong on Hacker News than this post.
- jjav 3mo ago> So pre-digital, no books, publications, photographic prints, scrolls, tablets, or clay etchings we lost to time? That is already answered in prior post. > Digital records are far far more long-lived — in the statistical average — than physical. Obviously that can't be known, simply because the data does not exist yet. There does not exist any digital data older than a hundred years, let alone a thousand. Will anything digital survive that long? Well, we don't know, come back in a thousand years to see. So far, the odds are stacked against it. Try to find just some tapes from the 70s to read, and you won't find much. And that's just ~50 years, a blink of an eye in terms of history. Meanwhile my magazine collection and photographs from the 70s are mostly as good as new, as usable today as back then.
- userbinator 3mo agoHistory is not only what "official" resources want you to believe. Those who want a "right to be forgotten" are really advocating for the "right to rewrite history".
- vbernat 3mo agoI have also restored the very first articles I wrote in the 90s when I was young <https://vincent.bernat.ch/en/blog/2026-old-web-articles https://vincent.bernat.ch/en/blog/2026-old-web-articles>. At the time, they were not "great," but now I think they have some limited historical values. For example, one of them is about how the national phone operator was billing minutes. The information was easy to find in the past but is pretty scarce now. I didn't use the wayback machine because it didn't archive everything I needed and because I still had the files on my hard drive, but if I didn't, I would have been happy to recover them.
- phendrenad2 3mo agoI agree, indiscriminate archival is doomed to fail. Spending all of your resources backing things up leaves you with no resources to curate. This is where archive.org goes wrong I think. I think identifying information that is valuable and distilling, reformatting, and republishing it is much more important.
- LegionMammal978 3mo agoThe problem is, some information is interesting only to a tiny audience long after the fact, so there's no chance anyone would've thought to curate it in the moment. E.g., for a few of my projects I like to dig up old versions of niche libraries to observe how their behavior has changed. But for most traditional distribution methods (putting the tarball up on your website or on a SourceForge-style website), the common assumption is that no one wants old broken versions, so you can regularly wipe everything but the last 1-3 versions to reduce clutter: the consequence is that the old versions accidentally saved on archive.org may be the only copies left anywhere. The more recent practice of publishing Git repos has considerably improved this situation.
- timbit42 3mo ago"Those who cannot remember the past are condemned to repeat it" - George Santayana
- zrn900 3mo ago> To what end, other than "well, future historians may need it"? Future historians ! will ! need it. Today we learn the most about past civilizations from what the ordinary people left behind.