6 ms·
Looks very nice. Relatedly, I wish I could automatically freeze and archive every single web page I visit, minus heavy media, possibly with very low quality im
by rjeli 6y ago
Looks very nice.
Relatedly, I wish I could automatically freeze and archive every single web page I visit, minus heavy media, possibly with very low quality images. I tried squid and the internet archive’s proxies, but MITM’ing myself is just slightly too annoying. There’s SingleFile[0] which does pretty much exactly what I want, ripping every single page into a self-extracting HTML+zip file, but it runs inside the browser so it adds a little delay after you navigate to a page, again slightly too annoying. Anyone have a recommendation for a seamless way to do this? Otherwise I’ll probably roll my own extension that pipes every URL to a local process that rips in the background with e.g. selenium.
I wish there were a way to run fully privileged extensions in Firefox, i.e. in the browser context instead of the page...
https://addons.mozilla.org/en-US/firefox/addon/single-file/ https://addons.mozilla.org/en-US/firefox/addon/single-file/
- Arubis 6y agoIf you wanted to keep this off your local box, you could create a separate pinboard.in account, pay for their archival service, and script your browser to pin every URL you hit.
- sriku 6y agoZotero can store webpage snapshots. Not sufficient? I also sometimes print-to-pdf.
- totetsu 6y agoIts not to hard to stand up your own local network webdav server to sync zotero to either. https://docs.bytemark.co.uk/article/run-your-own-webdav-server-with-docker/ https://docs.bytemark.co.uk/article/run-your-own-webdav-serv...
- interurban 6y agoSaving every single page feels a little overwhelming to me, I open lots of pages looking for a piece of information or an answer to a question and many aren't relevant, many others are outright spam. That said, I use pinboard to save/bookmark links, and I paid for the archival account type which automatically stores the pages I save. There's a handy bookmarklet so saving a page is a one click operation.
- treszkai 6y agoHow do you deal with the eventuality of Pinboard going down? (Which will almost certainly happen sooner or later.) While the civilization doesn't depend on my data, I always like to have a backup, so paying for a service to store my web archive is barely more future-proof than saving links.
- canadianwriter 6y agoYou can export the archive and put it on your own harddrive if you want, that's what I do, so I have an offline backup of all sites I have bookmarked.
- toyg 6y agoOne thing to note is that pinboard’s retrieval is not instantaneous. It will save your page “eventually”, which sometimes means “never” because the scraper will get there too late. It happened to me quite a few times, which is part of the reason I’m not a subscriber anymore.
- stewbrew 6y agoAs an alternative, there is also Save Page WE which works well for me. It takes some time to save a page but I cannot remember a delay when opening the resulting HTML. I personally save pages as plain text though so that I can use grep etc. later on. Pages with many essential graphics I also save as PDF or epub.
- ilSignorCarlo 6y agowhat's your workflow for saving the pages as PDF/epub?
- stewbrew 6y agoI use the "Save as ePub" addon or simply "Print to file". I save the files in a directory with all the thematic subdirectories in place and then run a script to distribute the file to several places. With earlier versions of firefox I used the shelve addon, which unfortunately isn't supported anymore.
- kfrzcode 6y agoA bit of scripting and this: https://archivebox.io/ https://archivebox.io/ might fit the bill? Or this: https://github.com/Rhizome-Conifer/conifer https://github.com/Rhizome-Conifer/conifer Maybe this: https://github.com/webrecorder/pywb https://github.com/webrecorder/pywb Or WWWOFFLE: http://www.gedanken.org.uk/software/wwwoffle/ http://www.gedanken.org.uk/software/wwwoffle/ Also pinboard.in but that's $
- te_ch 6y agoThanks! I haven't thought about storing the web pages. If I may, why would you like to keep a copy of the webpage you bookmark?
- maxerickson 6y agoBecause it's likely to vanish.
- danielheath 6y agoHeadless chrome can generate PDF files - perhaps you could use that for archiving?
- te_ch 6y agoI'll explore this one, thanks!
- frabbit 6y agoObviously not answering for the OP, but: because sites disappear or change and also having a local database that does not depend on internet access is wonderful. (Zotero is also great in this regard).
- te_ch 6y agoGot it. I just wonder... Storing web pages has a cost. Having notes for each bookmark (which I usually use to record key data points I find in web pages), and also notes at the topic and board levels (which can be used for similar purposes): is there still any benefit (above cost) of keeping an exact copy of the webpage you bookmarked? I guess it all depends on what's the chance of needing to go back to older bookmarks to re-visit something, right? Thanks for your feedback!
- luckydata 6y agoAnother thought: google has become very aggressive at indexing new content, so much so that old stuff becomes hard to find. There’s pages I know I visited that I just can’t find anymore, no matter how hard I use advanced search.
- corytheboyd 6y agoMaybe you could get clever with a service worker injected by an add-in. The worker context doesn’t have DOM access but maybe you could stream the DOM in chunks through postMessage and do some assembly in the worker. Though honestly that is a lot of hackery just to have it happen in the browser, but it does sound like an interesting experiment! Maybe I’ll try it out myself later..
- knrz 6y agoThe Chrome pageCapture API lets you dump the page as mhtml. It’s supposed to cache all content offline.
- SanchoPanda 6y agoSingleFile has a cli version which I use and like.
- numpad0 6y agoBack when I was lurking Futaba Channel, the standard protocol to share a stale post(they don’t archive posts) was to take MHT files through “Save As...” and uploading it. I think even some mobile apps for browsing Futaba had dedicated save/load features. Saves everything, even some ads.
- DenisM 6y agoCopy paste into email, send to yourself.
- ohxh 6y agoI'm working on something similar. In chrome you can request an image of the active tab from a background page, write it to a canvas to downscale it, and then dump it into indexeddb as a blob of base64. This only captures tabs that have been active at least once but I haven't noticed any additional latency using it.
- swsieber 6y agoIf you really wanted to you could roll your own browser with an Electron core - you'd be able to inject stuff there willy-nilly. At work we instrument our puppeteer (headless chrome scripted) tests so that it basically sumps out the Dom and other important rendering things. I then wrote a tool that lets one step through the dump with an interactive timeline showing the current state of the page along with console events, user generated events, etc. So the power is there if you want to get your hands really, really dirty.
- rjeli 6y agoYeah, I also looked into qutebrowser and Falkon, which let you run arbitrary python against the QtWebEngine bindings. That’s probably the best route, but neither of them fully support WebExtensions / have the ecosystem of chrome or Firefox.
- adamcanady 6y agoI recently did this using Electron's BrowserView. I had to step away from using electron because I encountered segfaults like `Received signal 11 SEGV_MAPERR 000000000060` just on visiting cnn.com and clicking one of their nav links (without much going on otherwise, which I found kinda crazy). Electron also only recently added support for PDFs, and it's still a little buggy. I was able to get around most of these with pdf.js. I ended up using a combination of chrome extension for injecting stuff into webpage + rpc to a local node.js server running inside an electron app (for convenience). I only made this for personal use, so I'm ok with <arbitrary_caveat>s.
- WrtCdEvrydy 6y ago> because I encountered segfaults like `Received signal 11 SEGV_MAPERR 000000000060` just on visiting cnn.com and clicking one of their nav links (without much going on otherwise, which I found kinda crazy). Yeah, web browsers have so much compatibility code this does not surprise me...
- lepht 6y agopinboard.in is a fantastic bookmarking service that will archive your bookmarked sites for you if you pay for the $39/year archival account: http://pinboard.in/tour/ http://pinboard.in/tour/
- laurentb 6y agoI think Worldbrain Memex does this: https://getmemex.com/ https://getmemex.com/ Been using it for a little while and it's definitely interesting although I'm not sure it backs up images as well, but could be worth a try for you.
- gildas 6y agoAuthor of SingleFile here. I think you could fix the delay issue by disabling some options (search "CPU consumption" in the help page of SingleFile). I might add an option in the extension for delegating to a "daemon" (based on the CLI code [1]) the capture of a page. [1] https://github.com/gildas-lormeau/SingleFile/tree/master/cli https://github.com/gildas-lormeau/SingleFile/tree/master/cli
- rjeli 6y agoBullseye! That's perfect. I should've spent more time playing around with the preferences. After browsing around a bit, the only page I notice a lag on is google search results, it freezes for ~300 ms after load. I'll keep messing around to see if I can eliminate that. Now just have to write a little bash script that sits in the background and moves the pages out of my download folder. (edit 5 seconds later): never mind, you can change the filename to put them in a subfolder. Thanks for your extension!
- dtakuma 6y agoYou my find something that fits you needs in the projects mentioned here https://github.com/pirate/ArchiveBox/wiki/Web-Archiving-Community https://github.com/pirate/ArchiveBox/wiki/Web-Archiving-Comm...
- eitland 6y agoI use Joplin with the web clipper extension. It isn't perfect but it is really really nice when it works.
- eg312 6y agoI made an extension to save web pages as eBooks: https://github.com/alexadam/save-as-ebook https://github.com/alexadam/save-as-ebook. You can easily modify it to remove the images (or scale them down) and invoke it automatically on each page visit or refresh.
- fouc 6y agoIt would be nice if we had a proper proxy that could run as a local daemon and basically caught every file request the browser made and saved the appropriate ones, possibly downsampling the images.