7 ms·
But how certain is the future of WayBackMachine, when disaster strikes, all your links are dead. On the other hand, the original links can still be read from th
by cornedor 6y ago
But how certain is the future of WayBackMachine, when disaster strikes, all your links are dead. On the other hand, the original links can still be read from the url, so the original reference is not completely gone.
- oblio 6y agoDoesn't the link to the WayBackMachine contain the original link?
- INTPenis 6y agoYeah, my thoughts were more of the way Waybackmachine is funded. I don't feel comfortable sending a bunch of web traffic to them for no reason other than it being convenient. The wayback machine is a web archival project, not your personal content proxy to make sure your links don't go stale. They need our help both in funding and in action, one simple action is not to abuse their service.
- sanitycheck 6y agoPrecisely my first thoughts, too. It's an archive, not a free CDN. I hope the author of this piece considers donating and promoting donation to their readers: https://archive.org/donate/ https://archive.org/donate/
- Lex-2008 6y agoWayBackMachine alternative, archive.is, has an option to download zip archive of HTML with images and CSS (but no JS) - this way you can preserve and host a copy of original webpage on your own website
- moonchild 6y agoOr just wget -rk... Mirroring a website isn't so hard that you need a service to do it for you. Your browser even has such a function; try ctrl-s.
- abricot 6y agoThe "SingleFile" plugin is a better version of ctrl+s. It will save all pages as single html file and even include images as an octet stream in the file so they aren't missed.
- peq 6y agoI would be careful in mirroring a site. It's very likely to violate copyright or similar laws, depending on where you are. I think archive.org is considered fair use, but if you put it on a personal or even business page it might be different. For example Google News in EU is very limited in what content they may steal from other web pages.
- dredmorbius 6y agoINTERNETARCHIVE.BAK: The INTERNETARCHIVE.BAK project (also known as IA.BAK or IABAK) is a combined experiment and research project to back up the Internet Archive's data stores, utilizing zero infrastructure of the Archive itself (save for bandwidth used in download) and, along the way, gain real-world knowledge of what issues and considerations are involved with such a project. Started in April 2015, the project already has dozens of contributors and partners, and has resulted in a fairly robust environment backing up terabytes of the Archive in multiple locations around the world. https://www.archiveteam.org/index.php?title=INTERNETARCHIVE.BAK https://www.archiveteam.org/index.php?title=INTERNETARCHIVE.... Snapshots from 2002 and 2006 are preserved in Alexandria, Egypt. I hope there's good fire suppression. https://www.bibalex.org/isis/frontend/archive/archive_web.aspx https://www.bibalex.org/isis/frontend/archive/archive_web.as...
- phendrenad2 6y agoI wish there were a way to get a low-rez copy of their entire archive. So, only text, no images, binaries, PDFs (other than PDFs converted to text which they seem to do). As it stands the archive is so huge, the barrier to mirroring is high.
- dredmorbius 6y agoAgreed. When scoping out the size of Google+, one of ArchiveTeam's recent projects, it emerged that the typical size of a post was roughly 120 bytes, but total page weight a minimum of 1 MB, for a 1% payload to throw-weight ratio. This seems typical of much the modern Web. And that excludes external assets: images, JS, CSS, etc. If just the source text and sufficient metadata were preserved, all of G+ would be startlingly small -- on the order of 100 GB I believe. Yes, posts could be longer (I wrote some large ones), and images (associated with about 30% of posts by my estimate) blew things up a lot. But the scary thing is actually how little content there really was. And while G+ certainly had a "ghost town" image (which I somewhat helped define), it wasn't tiny --- there were plausibly 100 - 300 million users with substantial activity. But IA's WBM has a goal and policy of preserving the Web as it manifests, which means one hell of a lot of cruft and bloat. As you note, increasingly a liability.
- nikisweeting 6y agoSo archive your links yourself with one of the many local-web-archiving tools. https://webrecorder.io https://webrecorder.io https://github.com/pirate/ArchiveBox https://github.com/pirate/ArchiveBox https://github.com/pirate/ArchiveBox/wiki/Web-Archiving-Community#other-archivebox-alternatives https://github.com/pirate/ArchiveBox/wiki/Web-Archiving-Comm...