12 ms·
Monolith – CLI tool for saving complete web pages as a single HTML file
- keyle 3y agoI am really loving these 'new' pure rust tools that are super fast and efficient, with lovely API/doco. Ah, it feels like the 90s again... Minus 50% bugs probably.
- snshn 3y agoHey, at least no memory leaks this time! Ü
- stringtoint 3y agoNice! Reminds me of the time I was working on a browser extension to do this.
- lagt_t 3y agoI remember IE5 was able to do this lol. It fell out of vogue for some reason, glad to see the concept is still alive.
- berkes 3y agoFirefox can still do it.
- thrdbndndn 3y agoChrome can too
- Hamuko 3y agoCan it? I'm only having Firefox save a bunch of files.
- phrz 3y agoSafari does this with .webarchive files
- toomuchtodo 3y agoRelated: Show HN: CLI tool for saving web pages as a single file - https://news.ycombinator.com/item?id=20774322 https://news.ycombinator.com/item?id=20774322 - August 2019 (209 comments)
- joeyhage 3y agoIt would be awesome to see support for following links to a specified depth, similar to [Httrack](https://www.httrack.com/ https://www.httrack.com/)
- codetrotter 3y agoI made a basic crawler using Firefox, thirtyfour https://docs.rs/thirtyfour/latest/thirtyfour/ https://docs.rs/thirtyfour/latest/thirtyfour/ and squid Basically, I took a start URL for the crawl, and my program would load the page in Firefox using thirtyfour, and then extract all links from the page and use some basic rules for keeping track of which ones to visit and in which order. I had Squid proxy configured to save all traffic that passed through it. It worked ok-ish. I only really stopped that project because of a hardware malfunction. The main annoyance that I didn’t get around to solving was being more smart about not trying to load non-html content that was already loaded anyway as part of the page. Because the way I extracted links from the page I also extracted URLs of JS, CSS etc that were referenced.
- gildas 3y agoYou can have a look at the last 2 examples here [1]. [1] https://github.com/gildas-lormeau/single-file-cli?tab=readme-ov-file#run https://github.com/gildas-lormeau/single-file-cli?tab=readme...
- arp242 3y agoI wrote something very similar a few years ago – https://github.com/arp242/singlepage https://github.com/arp242/singlepage I mostly use it for a few Go programs where I generate HTML; I can "just" use links to external stylesheets and JavaScript because that's more convenient to work with, and then process it to produce a single HTML file.
- lopkeny12ko 3y agoHow does this compare to SingleFile? https://www.npmjs.com/package/single-file-cli https://www.npmjs.com/package/single-file-cli
- gildas 3y agoAuthor of SingleFile here, one of the major differences is that monolith doesn't use a web browser to take page captures. As a result, it doesn't support JavaScript, for example. SingleFile, on the other hand, requires a Chromium-based browser to be installed. It should also produce smaller pages and is capable of generating ZIP or self-extracting ZIP files. However, it will take longer to capture a page. Note that since version 2, it is now possible to download executable files of the CLI tool [1]. [1] https://github.com/gildas-lormeau/single-file-cli/releases https://github.com/gildas-lormeau/single-file-cli/releases
- tiagod 3y agoI've been using SingleFile for ages now... it's my favorite browser extension after uBlock, thank you for your great tool! :)
- darkteflon 3y agoSingleFile is amazing - use it tens of times every day across desktop and mobile. Can’t recall a single instance of it breaking. Thank you sincerely for your excellent work.
- gildas 3y agoThanks a lot! Believe me, there have been a lot of bugs (+900 issues closed today) because it's hard to save a web page actually. You were lucky not to suffer ;)
- darkteflon 3y agoI bet! The proof of that must surely be in how poor a job formats like .webarchive do of it. SingleFile just makes this one really complex, really important thing trivially easy, and in a portable format. For anyone curating a knowledge base it’s an absolute godsend. I didn’t see any donation instructions on your GitHub - I for one would certainly love to chip in if you could point me in the right direction?
- simonw 3y agoWell this is fun... from the README here I learned I can do this on macOS: /Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \ --headless --incognito --dump-dom https://github.com > /tmp/github.html And get an HTML file for a page after the JavaScript has been executed. Wrote up a TIL about this with more details: https://til.simonwillison.net/chrome/headless https://til.simonwillison.net/chrome/headless My own https://shot-scraper.datasette.io/ https://shot-scraper.datasette.io/ tool (which uses headless Playwright Chromium under the hood) has a command for this too: shot-scraper html https://github.com/ > /tmp/github.html But it's neat that you can do it with just Google Chrome installed and nothing else.
- samstave 3y agoYay! I love Shot Scrapeer - I wish you had made it a decade ago! Thanks for shot scraper. Off the top of you head what would be the easiest command to have shotscraper barf a directory of shot-scraper HTMLs each day from my daily browsing history. This would be interesting if I have a browsing session for learning something and I am researching across a bunch of sites - roll it all up into a Digi-ography of the sites used in learning that topic? --- I've always been baffled that this isnt an inate functionality in any app/OS - its a damn computeer - I should have a great ability to recall what it displays and what you have been doing. Heck - we need our machines to write us a daily status report for what we did at the end of each day. Surely that would change productivity. If you were force to do a self-digital-confession and stare you ADHD and procrastination right in the face.
- jimmySixDOF 3y agolook at ArchiveBox from the comments below
- genewitch 3y agoThis used to be fairly simple to do before https everywhere, just install squid (or whatever) and cron the cache folder to a zip file once a day or whatever. There's paid solutions that kinda do what you want, but they capture all text on your screen and OCR it to make it searchable, which at least lets you backtrack and has the added advantage that it will make pdfs, meme images, etc searchable, too. last i heard it was mac only but a few folks mentioned some windows software that does it too. as an aside i don't consider reading/learning nearly all day to be a net negative, even if ADD is to blame. (i haven't had the "H" since i was a child.) A status report wouldn't "stare" me in the face; in fact, it would be nice to have some language model take the daily report and over time suggest other things to read or possible contradictions to link to.
- al_borland 3y agoI use read-it-later type services a lot, and save more than I read. On many occasions I've gone back to finally read things and find that the pages no longer exist. I'm thinking moving to some kind of offline archival version would be a better option.
- amcpu 3y agoI use a locally hosted YaCy instance with cached results to work around this scenario. Much of the content I am interested in is kept locally, so it’s good enough. When I have a bunch of “read later” tabs that pile up, I copy all their URLs into the crawler form with “Store to Web Cache” checked and it accomplishes what I described. Just another option to consider.
- nelsonfigueroa 3y agoI've used ArchiveBox in the past and it's been great for this purpose: https://github.com/ArchiveBox/ArchiveBox https://github.com/ArchiveBox/ArchiveBox
- saganus 3y agoThis is great, thanks!
- hu3 3y agoHi! This seems amazing and sustainable since it leverages industry standard tools such as yt-dl and chrome headless. Now I'm curious, what made you stop using it?
- nelsonfigueroa 3y agoI just found myself archiving fewer things over time and it’s been a while since I’ve saved anything. There’s nothing wrong with it though. In fact, I still have it on my machine.
- arp242 3y agoI have a lot of old unsorted bookmarks of "I want to look in to this, but don't have time now". Newer stuff is more organized, but I exported the old stuff and haven't looked at them in about five years. Last week I started organizing them a bit, and it's shocking how much is a 404. Even from major newspapers and such. I have no idea why anyone would take down old content (outside of some specific and rare reasons). Some are also on neither internet archive or archive.today.
- ethanpil 3y agoNice. My next step: Figure out how to make a web extension 1 click button. Tab to Monolith to Joplin with a tag.
- gildas 3y agoYou could download SingleFile [1], configure a WebDAV server in the options page (cf. "Destination" section), and set up Joplin to synchronize with the server. [1] https://github.com/gildas-lormeau/SingleFile https://github.com/gildas-lormeau/SingleFile
- dohello1 3y agoand I thought my code pages were long haha
- causality0 3y agoHow's this better than the MHTML functionality built into my browser?
- gildas 3y agoYou can find a comparison of file formats here: https://github.com/gildas-lormeau/SingleFile?tab=readme-ov-file#file-format-comparison https://github.com/gildas-lormeau/SingleFile?tab=readme-ov-f...
- russellbeattie 3y agoIf anyone is interested, I wrote a long blog post where I analyzed all the various ways of saving HTML pages into a single file, starting back in the 90s. It'll answer a lot of questions asked in this thread (MHTML, SingleFile, web archive, etc.) https://www.russellbeattie.com/notes/posts/the-decades-long-html-bundle-quagmire.html https://www.russellbeattie.com/notes/posts/the-decades-long-...
- rnewme 3y agoCool post. You should make hn entry
- andai 3y agoI always ship single file pages whenever possible. My original reasoning for this was that you should be able to press view source and see everything. (It follows that pages should be reasonably small and readable.) An unexpected side effect is that they are self contained. You can download pages, drag them onto a browser to use them offline, or reupload them. I used to author the whole HTML file at once, but lately I am fond of TypeScript, and made a simple build system to let me write games in TS and have them built to one HTML file. (The sprites are base64 encoded.) On that note, it seems (there is a proposal) that browsers will eventually get support for TypeScript syntax, at which point I won't need a compiler / build step anymore. (Sadly they won't do type checking, but hey... baby steps!)
- phatskat 3y agoDo you happen to have any resources on writing games in TS? I will def google, but game development has always been oddly hard for me and I’ve been in typescript land for a while now so figure it’s a good time to try again
- toastedwedge 3y agoIf I may ask, where can I read more about this? I wouldn't really know where to look for something like that, I'm afraid. Edit: wording.
- teaearlgraycold 3y agoThe proposal: https://github.com/tc39/proposal-type-annotations https://github.com/tc39/proposal-type-annotations
- slmjkdbtl 3y agoI used to only do single file HTML pages too, until I have a page that have multiple occurrences of the same image, it's wasteful to have the dataurl string every time that img occurs. Maybe I can save the dataurl string in JS once and assign it to those img in JS, but most of the time my page doesn't have any JS it feels bad to use JS just for this.
- 3y ago
- andai 3y agoDoes anyone know how an entire website can be restored from Wayback Machine? A beloved website of mine had its database deleted. Everything's on Internet Archive, but I think I'd have to (1) scrape it manually (they don't seem to let you download an entire site?), (2) write some python magic to fix the css URLs etc so the site can be reuploaded (and maybe add .html to the URLs? Or just make everything a folder with index.html...) It seems like a fairly common use case but I barely found functional scrapers, let alone anything designed to restore the original content in a useful form.
- belthesar 3y agoI bet the ArchiveTeam might be able to help you out with this. They were quite helpful when I wanted to make sure a site was preserved, and might be able to help you as well, or at least point you in the right direction. https://wiki.archiveteam.org/ https://wiki.archiveteam.org/
- gildas 3y agoIt's documented here: https://wiki.archiveteam.org/index.php?title=Restoring https://wiki.archiveteam.org/index.php?title=Restoring
- AdieuToLogic 3y agoOr perhaps wget[0] as described here[1] and documented here[2] could do the trick. 0 - https://www.gnu.org/software/wget/ https://www.gnu.org/software/wget/ 1 - https://tinkerlog.dev/journal/downloading-a-webpage-and-all-of-its-assets-with-wget https://tinkerlog.dev/journal/downloading-a-webpage-and-all-... 2 - https://www.gnu.org/software/wget/manual/wget.html https://www.gnu.org/software/wget/manual/wget.html
- mattsan 3y agoThis is addressed in the README and a comparison is given
- AdieuToLogic 3y ago> This is addressed in the README and a comparison is given The only mention of wget in the README reads thusly: If compared to saving websites with wget -mpk, this tool embeds all assets as data URLs and therefore lets browsers render the saved page exactly the way it was on the Internet, even when no network connection is available. This is not the only way to invoke wget in order to download a web page along with its assets. Should the introduction article I referenced above be deemed insufficient, consider this[0] as well. 0 - https://simpleit.rocks/linux/how-to-download-a-website-with-wget-the-right-way/ https://simpleit.rocks/linux/how-to-download-a-website-with-...
- dosourcenotcode 3y agoA cool tool to be sure. However I feel this tool is a crutch for the stupid way browsers handle web pages and shouldn't be necessary in a sane world. Instead of the bullshit browsers do where they save a page as "blah.html" file + "blah_files" folder they should instead wrap both in folder that can then later be moved/copied as one unit and still benefit from it's subcomponents being easily accessed / picked apart as desired.
- genewitch 3y ago"save as [single] html" or whatever hasn't worked reliably in over a decade. I wrote a snapshotter that i could post in a slack alternative "!screenshot <URL>" and it would respond (eventually) with an inline jpeg and a .png link of that URL. As i mentioned upthread, this worked for a couple of years (2017-2020 or so) and then it became unreliable on some sites as well. as an example, old.reddit.com hellthread pages would only render blank white after the first couple dozen comments. I haven't had the heart to try it with singlefile, but now that there's at least 3 tools that claim to do this correctly, i might try again. This tool, singlefile (which i already use but haven't tested on reddit yet) and archivebox. 4 tools, if you count the WARC stuff from archive.org
- jchook 3y agoHm, very interesting, especially for bookmarking/archiving. I'm curious, why not use the MHTML standard for this? - AFAIK data URIs have practical length limits that vary per browser. MHTML would enable bundling larger files such as video. - MHTML would avoid transforming meaningful relative URLs into opaque data URIs in the HTML attributes. - MHTML is supported by most major browsers in some way (either natively in Chrome or with an extension in Safari, etc). - MIME defines a standard for putting pure binary data into document parts, so it could avoid the 33% size inflation from base64 encoding. That said, I do not know if the `binary` Content-Transfer-Encoding is widely supported.
- snshn 3y agoMHTML support is planned, there's a couple of other problems that need to be resolved first, but it's a good format for archiving, been requested many times
- jchook 3y agoThanks for the reply. Very exciting. I would love to see MHTML support on this.
- Hamuko 3y ago>MHTML is supported by most major browsers in some way Firefox? What about mobile versions of browsers?
- jchook 3y agoWe should submit a PR
- sunshine202022 3y agofun
- publius_0xf3 3y agoAwesome tool. A note to the devs: the latest version on winget is v2.7.0, which is several months behind the latest version.
- fagrobot 3y agohttps://github.com/gildas-lormeau/SingleFile https://github.com/gildas-lormeau/SingleFile
- victorbjorklund 3y agoThis is great. I have wished for something like this.
- max_ 3y agoIt still blows my mind that browsers don't provide features this out of the box.
- Alifatisk 3y agoI think they do? Have you tried hitting cmd+s or ctrl+s? You can save webpages like that. But I don’t know if they can compress everything into a single html file though.
- max_ 3y agoAlot of the CSS & JavaScript is usually broken with ctrl+s. A great option used to be the mhtml format chrome. (It had to be enabled in chrome flags) But mhtml seemed to be removed from chrome since recently.
- vanderZwan 3y agoLast time I tried that it saved a static version of the current DOM, instead of the page source. I'm assuming that the reasoning behind that is that most people want to save a snapshot of what they are currently seeing, and that this is the easiest way to have somewhat reliable results for that.
- snshn 3y agoSo true. Monolith is using libraries made by Mozilla for their Rust-driven browser engine (which I believe, never happened to be). I really would love for it to be a part of some browser one day, the demand is clearly there. Nobody likes to have a file+folder abomination on their drive, or some shady formats like .webarchive
- Gormo 3y agoThe MHTML format [1] has been around for 25 years and was natively supported by multiple browsers for decades. Modern browsers have regressed in functionality. [1]: https://en.wikipedia.org/wiki/MHTML https://en.wikipedia.org/wiki/MHTML
- hu3 3y agoChrome does support it: https://i.imgur.com/HF7GXEI.png https://i.imgur.com/HF7GXEI.png
- k1ck4ss 3y agoHow would I archive an on-prem hosted redmine solution (https://www.redmine.org/ https://www.redmine.org/)? It is many, many years old and I want to abandon it for good but save everything and archive it. Is that possible with monolith?
- planb 3y agoYou're probably better off with a recursive wget here. IIRC redmine was not really javascript heavy and monolith looks to me like it only saves one page.
- farzadmf 3y agoIronically, I decided to try with the repo's own Github page, and when I open the resulting HTML file in Chrome, it's all errors in the console, and I don't see the `README` or anything
- yencabulator 3y agoGithub is a pile of Javascript that adds things to the DOM browser-side, the monolith README specifically says it does not run Javascript, and shows you a workaround for when that matters.
- AdmiralAsshat 3y agoSo what happens if the page is behind a paywall and the embedded Javascript stores some authentication or phone-home code? Does that end up getting invoked on the monolith copy HTML? I'm wondering how this would work if I wanted to use it to, say, save a quiz from Udemy for offline review.
- hollander 3y agoyt-dlp can use browser cookies to get access to Facebook videos. This should have something similar. yt-dlp --cookies-from-browser firefox https://www.facebook.com/1234videos/5678/
- katejmurchison 3y ago[dead]
- fs111 3y agohttps://en.wikipedia.org/wiki/WARC_(file_format) https://en.wikipedia.org/wiki/WARC_(file_format)
- pbnjeh 3y agoDoes anyone remember the Firefox extension Scrapbook, from "back in the day"? I used to use it a lot. Look "back" 5 - 10 years, or more, and it's striking how many web resources are no longer available. A local copy is your only insurance. And even then, having it in an open, standards compliant format is important (e.g. a file you can load into a browser -- I guess either a current browser or a containerized/emulated one from the era of the archived resource). Something that concerns me about JavaScript-ed resources and the like. Potentially unlimited complexity making local copies more challenging and perhaps untenable.