6 ms·
It saves pages to archive.org as well. You might want to be careful while using this to archive personal content.
by weekay 5y ago
It saves pages to archive.org as well.
You might want to be careful while using this to archive personal content.
- kenniskrag 5y agoCan be disabled but on in default mode. https://github.com/ArchiveBox/ArchiveBox/wiki/Configuration#submit_archive_dot_org https://github.com/ArchiveBox/ArchiveBox/wiki/Configuration#...
- Proven 5y ago"Democratizing Tragedy of Commons at scale"
- asaddhamani 5y agoarchive.org will only archive publicly visible content and it respects robots.txt
- tyingq 5y ago"and it respects robots.txt" Not since 2017. https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea... They now have a clunky manual process to exclude your site. https://help.archive.org/hc/en-us/articles/360004651732-Using-The-Wayback-Machine https://help.archive.org/hc/en-us/articles/360004651732-Usin... ("How can I exclude...") They don't spoof user agents, but blocking them actively doesn't remove their history.
- Karunamon 5y agoWhen did this change? It used to be that adding robots.txt would retroactively remove archives for a domain.
- tyingq 5y ago2017.
- pseudalopex 5y agoHide. Not remove.
- nikisweeting 5y agoYes, I can write a long article about why it's the default someday. I've agonized over this decision for many many months, and it's flipped flopped a few times as well. The short version is that defaults in software are really important (90% of users wont change them), and I don't trust myself to code ArchiveBox 100% correctly so as to never lose data, or the majority of people to store their archives correctly so as to never lose data on their own. Archive.org is the redundant failsafe. Another good reason is that Archive.org is not the only way that your archive content can be leaked, the security model means that archived pages can read each other's content, so I want to make it abundantly clear to users that by default it's designed to only archive content thats already public (in which case it's already fair game for Archive.org). I've settled on leaving it on as the default, but I do mention 3 times in the README how to disable it, most notably in the CAVEATS section which explains both the security model drawbacks and how to prevent your content from being leaked to Archive.org or other 3rd party APIs.
- jka 5y agoAlthough I tend privacy-by-default for most deployed technologies, the context of archiving does change the criteria quite a lot; you've selected a sensible and reasonable default, I reckon. Hopefully integrity is a consideration too? Glad to read that article, one day :)
- nikisweeting 5y agoIntegrity is absolutely paramount too of course, which is why I chose Django (because of the mature DB migrations system that makes upgrades deterministic, reversible, and relatively painless). Hand coding a schema migration system would be a recipe for disaster and an easy opportunity for users to lose data. Nevertheless, no system is perfect, and even with Django helping guard database integrity and multiple redundant index files, it's possible I'll make a mistake someday that leads to data loss on upgrade. I don't want that situation to be the next (mini) library of Alexandria, and saving copies to Archive.org helps serve as a last-resort backup.
- 22c 5y agoAppreciate the project, web archives are becoming an important part of the internet ecosystem. I personally have no issue with the defaults, but if you've agonized over the defaults, perhaps you should consider clearly documenting it in the main project README instead of leaving it for people to find in the config documentation.