4 ms·
I am not angelmm, but another happy ArchiveBox user. My choice of extractors is the following: Singlefile, PDF, Screenshot, archive.org. I found the largest i
by Linux-Fan 3y ago
I am not angelmm, but another happy ArchiveBox user.
My choice of extractors is the following: Singlefile, PDF, Screenshot, archive.org.
I found the largest issue with any website archiving tools to be the discrepancy between what I see in my Web Browser and what is saved. The most "reliable" way that still works today for me seems to be the "Save Page WE" Firefox plugin.
I have a sidecar container running that checks for HTML files appearing in a directory, triggers the archivebox save and then overwrites the "singlefile" capture by the provided HTML file. This way, I can trigger archiving by just using the Save Page WE plugin and storing the resulting HTML file in the directory.
- zerkten 3y agoCan archive box aggregate content "bookmarked" in different places? I want a tool that will pull saved items on Reddit, favorite posts on HN, etc. in addition to bookmarks posted to pinboard.in to a single place. In many ways, the share functionality in iOS allows me to get all URLs into a single place, but this doesn't help on desktop. I know this isn't necessarily and easy task, if APIs aren't available. I'd be OK with a client component or browser extension, if it was open source and self-hostable.
- Linux-Fan 3y agoAFAIK it does not contain any function to support this directly. My primary way to interact with Archive Box for such purposes is by calling it on the command line. Scripts may be used to obtain the URLs of interest from any source. When I started with Archive Box I had some existing downloaded Websites from ScrapBook and Save Page WE already. I used some hacky scripts to extract the URLs from the respective pages and overwrite Archive Box' downloaded copies by my original copies as to make it work for pages that had been deleted in the meantime. All my data sources were local/desktop though.
- nikisweeting 3y agoYeah, you can set up scheduled imports from browser history, bookmarks services, RSS feeds, etc. There are instructions for most of the common sources in the Input Sources section of the ArchiveBox readme :)