4 ms·
I've been meaning to set up some kind of basic crawl & archive system forever. Ideally I'd like to output something replayable and analyzable like a WARC or HAR
by PowerfulWizard 7y ago
I've been meaning to set up some kind of basic crawl & archive system forever. Ideally I'd like to output something replayable and analyzable like a WARC or HAR, but also spit out a PDF, which I think chrome headless should be able to do. Right now I just print to file if I want to "save" something. But your write-up is very thorough, that is basically the situation.
So much good unique content on youtube too which is almost impossible to properly archive due to size, my subscriptions alone would probably be over 1TB.
I'd like to write a basic tool that would take a PDF, i.e. of a book, and output a directory of PDFs of snapshots of all web links in the book, to create basically a full reference snapshot for a given book that could be stored alongside the book. Not sure if it work well and result in a reasonable size.
- gwern 7y ago> Ideally I'd like to output something replayable and analyzable like a WARC or HAR, but also spit out a PDF, which I think chrome headless should be able to do. ArchiveBox does WARCs and PDFs, and does embedded media; it's easy to use, you can point it at a newline-delimited textfile of URLs and it'll process it. I'm not sure how it handles YouTube - whether it shells out to something like youtube-dl or not... But really, 1TB is not all that much. You can get 8TB internal HDDs for like $200 now.