4 ms·
gwern has a very involved post on archiving as well, https://www.gwern.net/Archiving-URLs https://www.gwern.net/Archiving-URLs Somewhere on my to-do list is ar
by patrickyeon 8y ago
gwern has a very involved post on archiving as well, https://www.gwern.net/Archiving-URLs https://www.gwern.net/Archiving-URLs
Somewhere on my to-do list is archiving everything I visit on the internet. It's frustrating to know that I've seen something, but be unable to find it again.
- gildas 8y agoFor this, you could use SingleFile [1], an extension for Chrome and Firefox. It can auto-save pages. [1] https://github.com/gildas-lormeau/SingleFile https://github.com/gildas-lormeau/SingleFile
- vackosar 8y agounfortunately it requires a lot of permissions. I try to minimize exts like this
- toomuchtodo 8y agoI use grab site for extensive archiving operations. It’s not an extension, but trivial to launch from a command line (and I stuff the data into a Backblaze B2 account for later Internet Archive ingestion). You could take this and post process your internet history on a rolling basis to accomplish your goal. https://github.com/ludios/grab-site https://github.com/ludios/grab-site
- gildas 8y agoI really did my best to minimize the APIs used by the extension. Note that Chrome 70 allows you to restrict extensions by host [1]. [1] https://blog.chromium.org/2018/10/trustworthy-chrome-extensions-by-default.html https://blog.chromium.org/2018/10/trustworthy-chrome-extensi...
- vackosar 8y agotool for domain to epub conversio https://github.com/haroldtreen/epub-press-clients/blob/master/README.md https://github.com/haroldtreen/epub-press-clients/blob/maste...
- pflanze 8y agoI should probably try this. I wonder how it will compare with what I'm doing: I simply use the browser's (Firefox) save page feature (ctl-s). Then every now and then, I convert the folder with these pages to a squashfs image (which de-duplicates all the CSS, JS, image files that are saved multiple times). I then use shell tools to search (ls grep locate etc.). This doesn't save the URL, but I also maintain a private "bookmarks" Git repository for more interesting bits where each bookmarked resource gets its own file (in a descriptive hierarchy), along with thoughts and notes. What works well is that I am somewhat selective, complete trash doesn't end up being archived, it's pretty space efficient. It's also simple to wade through (in the shell), each saving action is just a file and folder pair. Also often I use the reader mode Firefox feature, and then save that (i.e. I save what I saw, not what was delivered). What doesn't work so well is that saving the page is often a bit of a hassle, often Firefox reports save errors then I just ctl-s again to see if the previous attempt succeeded, if not, hit the reload icon in the download task list which apparently forces it. Also, that it saves embedded videos, haven't figured out yet how to disable that, so periodically I remove videos again. Also, when saving multiple pages from the same site, links to other pages don't go to the local mirror. (In cases where it's important, I use wget -m -k.)
- gildas 8y agoI think the differences would be the following ones: - You wouldn't have to bother with JS files, they are removed from saved pages by default. - You couldn't (easily) de-duplicate resources because they are embedded in base64 into the page. However, SingleFile can detect all the hidden elements and unused CSS rules/declarations (by computing the cascade). So the HTML and CSS are optimized. It is also able to group duplicate images (by using CSS custom properties). There are other options to make the document size as small as possible. Most of the time, pages saved with SingleFile are smaller than Chrome MHTML files. - The saving process would be much more simple, reliable and could be automatic. - Videos wouldn't be saved (by default), a snapshot of the video would replace each video. - Maybe the "filename template" in SingleFile would help you to organize things.
- victor106 8y agoSingleFile does not seem to embed images in the file. I tried this[1] and it seems to work fine. It seems to be closed source though..:( [1]https://chrome.google.com/webstore/detail/save-as-mht/hfmodljjaibbdndlikgagimhhodmobkc?hl=en https://chrome.google.com/webstore/detail/save-as-mht/hfmodl...
- gildas 8y agoIt does embed images (and all other resources). Can you post a URL showing this issue please?
- victor106 8y agoI used it on an internal website and when I opened it, it didn't have the image in it. In place of images I only see blank place holders. But it did preserve the structure of the website.
- gildas 8y agoIf you see JS errors related to the extension or HTTP errors in the console of the developer tools, please file an issue with the errors on GitHub. It would help understanding what's going wrong. EDIT: it could be related to the fact that SingleFile is maybe more strict regarding the HTTP header "Content-Type". For example, it will discard images with "text/html" as content type value.
- namibj 8y agoYou are not alone. It's further back on said to-do list, as it appears to be a problem with no easy existing solution, and it's too big of a project for the immediate payoff. If you happen to stumble upon a solution (obviously self-hosted/local, due to the unfiltered access to page content), I might be willing to contribute with configuration scripting or client/server splitting so that the bulk doesn't have to stay on e.g. a laptop.
- anarcat 8y agoI toyed with that idea a few times, but this now strikes me as a problem for two reasons: 1. it's a security liability. some content i load in a web browser is private and I don't want to archive or duplicate it anywhere 2. it means a lot of crap. part of archiving content is curating what gets archived and what doesn't. i didn't touch on this in the article, but it's a key idea archivists need to address. for example, archiving the page I'm typing in right now (news.y.com/reply) makes no sense at all because it's solely dynamic content and would mean nothing when browsed later. So instead I send specific links I want to keep in my bookmarks system, as I mentioned in the article. It's far from ideal, but it's a much better compromised than archiving every time I visit my weather service page. ;)