3 ms·
> Running it over my bookmarks > Once I’d written the initial version of this script and put all the pieces together, I used it to create webarchives for 6000
by tedmiston 2y ago
> Running it over my bookmarks
> Once I’d written the initial version of this script and put all the pieces together, I used it to create webarchives for 6000 or so bookmarks in my Pinboard account. It worked pretty well, and captured 85% of my bookmarks – the remaining 15% are broken due to link rot. I did a spot check of a few dozen archives that did get saved, and they all look good.
I was a tad confused by this part.
Did you (or how did you) verify that the headlessly saved web archives for thousands of bookmarks visually match the pages shown in the browser?
This is the biggest problem I've had with command-line archival tools: they save some version of the page, but it often differs substantially from what I actually see in my browser — things like pop-up artifacts covering the page or news articles are full of ads that are otherwise blocked in my headed browser.
The SingleFile extension for Chrome works more completely and accurately than anything else I've come across so far, but it does still break weirdly sometimes too.
I would love to find a programmatic way to automate the visual verification, e.g., archiving a page with multiple different tools and visually diffing the rendered pages across tools with small margins of error. Maybe someone else has worked on this already.
- kwhitefoot 2y agoWebScrapbook is also worth a look. I find that I like it slightly better than SingleFile for creating copies that are not packaged as single files. This lets me hard link identical asset files to save space.
- byteknight 2y agoArchivebox with its multi-method backups. Screenshots, text, html, etc.
- nikisweeting 2y agoI also built AI-based quality assurance automation to check archive quality for archivebox recently, it works surprisingly well.