4 ms·
A shame how many of these links are now pointing to 404 not found pages. People assume that once it's on the internet it's there forever, but it's really not t
by qqtt 5y ago
A shame how many of these links are now pointing to 404 not found pages.
People assume that once it's on the internet it's there forever, but it's really not true. Some of our favorite articles and insights from past days are long gone.
- MisterBiggs 5y agoAs someone who uses the search function on HN regularly I have to agree its pretty jarring to see how many webpages totally disappear. Thankfully the wayback machine exists.
- yen223 5y agoOne day, the wayback machine will start returning 404s too
- AnimalMuppet 5y agoThat's... a bit sobering. At that point, we will really have lost a lot of history.
- twodai 5y agoIt's possible some other group will buy the data, or they will make it easily download able
- jlokier 5y agoLike the way Google did with the Usenet archive. Make it available via a searchable interface, then gradually degrade the archive until old posts can't be found any more.
- squarefoot 5y agoThey also degraded the interface so that the posts that once would be found via normal searching now aren't anymore. After the removal of the discussion filter, searches for X started returning more sites selling X than forums talking about X. https://www.reddit.com/r/google/comments/2b54ux/google_completely_removed_discussions_search_is/ https://www.reddit.com/r/google/comments/2b54ux/google_compl... There was a thread on Google forums about that, and I recall many upset users asking for the option to be reactivated, but (the irony) that discussion was removed as well.
- metagame 5y agoActually, archive.org is already easily downloadable, in that you can already help host the dweb copy. archive.org, available via webtorrent: https://dweb.archive.org/ https://dweb.archive.org/ It's pretty slow, though.
- dotancohen 5y agoI think that the Wayback machine supports either an HTTP status code or meta tag that causes it to not serve previously-cached contents. Therefore I try to save webpages that I care about, but it's getting harder and harder. Not to mention the space it takes - is it really worth the hundreds of GiB of personal archives when I'll likely want - not need - maybe a few KiBs of it decades down the line. And even then, will I be able to find it? Sometimes I think about just archiving a screenshot and the text of websites, instead of markup and related files.
- dredmorbius 5y agoOne day there will be no more 404s or 200s.
- omn1 5y agoI maintain a couple of bigger Github repos and blogs and was shocked at how many links break on a regular basis, so I wrote my own link checker in Rust [1] and started donating to the Wayback Machine. Please consider doing so, too as the web would be a worse place without them. [2] [1]: https://github.com/lycheeverse/lychee https://github.com/lycheeverse/lychee [2]: https://archive.org/donate/ https://archive.org/donate/
- hanniabu 5y agoPersonally I figure that if my link happens to end up dead that anybody that really wanted to see it could check it out on web archive.
- fedbook 5y agoI see an opportunity here… in the days of multiterrabyte hard drives, why not make an extension that simply saves the HTML and images for every site you visit? Maybe not appropriate for a phone, but perhaps there could be an option to sync browser activity…
- vimy 5y agoI don’t remember the name but a project like that exists on github.
- sixstringtheory 5y agohttps://github.com/ArchiveBox/ArchiveBox https://github.com/ArchiveBox/ArchiveBox