15 ms·
Show HN: CLI tool for saving web pages as a single file
- mehrdadn 7y agoThis is awesome. One question though: how does it handle the same resource (e.g. image) appearing multiple times? Does it store multiple copies, potentially blowing up the file size? If not, how does it link to them in a single HTML file? If or if so, is there any way to get around it without using MHTML (or have you considered using MHTML in that case)? Also, side-question about Rust: how do I get rid of absolute file paths in the executable to avoid information leakage? I feel like I partially figured this out at some point, but I forget.
- chmod775 7y agohttps://github.com/Y2Z/monolith/blob/master/src/html.rs#L94 https://github.com/Y2Z/monolith/blob/master/src/html.rs#L94 Hope this answers your question (it gets converted to a data URI, and there's apparently no de-duplication).
- flatroze 7y agoThank you! It's pretty straight-forward: this program just retrieves assets and converts them into data-URLs (data:...), then replaces the original href/src attribute value, so in case with the same image being linked multiple times, monolith will for sure bloat the output with the same base64 data, correct. I haven't looked into MHTMTL, ashamed to admit it's the first time I'm hearing about that format. I need to do some research, maybe I could improve monolith to overcome issues related to file size, thank you for the tip! And about Rust: I think you're way ahead of me here as well, this is my first Rust program. If you're talking about it embedding some debug info into the binary which may include things like /home/dataflow then perhaps there's a compiler option for cargo or a way to strip the binary after it's compiled. ¯\_(ツ)_/¯ Sorry, that's the best I can tell at the moment.
- mehrdadn 7y agoOkay thanks! That was a pretty quick reply :) Regarding MHTML, it's basically the MIME format emails are in (which are basically inherently self-contained HTML documents). Various browsers have had varying degrees of support for it over the years. Chrome recently made it harder to save in MHTML format; I don't know how long they will be able to read the format, so I can't guarantee that if you go in that direction it'll still be useful for a long time, but at the moment there is still some support for it.
- dspillett 7y agoOne way to dedupe inline image resources while still using HTML rather than MHTML, could be to encode them in css once, and transform the image element to something with that class.
- mehrdadn 7y agoThat'd easily break Javascript though.
- dspillett 7y agoGood point. I was thinking in the direction of something I'm tinkering with in a similar area. There getting a static snapshot of the current DOM or fragment is key (meaning scripts being stripped out is an intentional feature). Tweaking the document contents for efficiency could significantly impact a lot of script work that may be present.
- gildas 7y agoSingleFile which can run on CLI too handles this by using CSS variables.
- ajxs 7y agoVery cool. Have you considered incorporating an option for following links within the same domain to a certain depth? I remember using tools such as this in the past to save all the content from certain websites.
- flatroze 7y agoThank you! I'll add it as an issue, since it could definitely be useful for "archiving" certain resources more than 1 level deep. Do you remember the name of that tool by any chance?
- xtreme 7y agoI have used HTTrack in the past for this.
- ajxs 7y agoI wish I could remember the exact tool, this was over a decade ago. If you just do a quick internet search for such a tool you'll likely find whatever I used, it certainly wasn't anything sophisticated. It was a Windows GUI tool designed specifically for the task. Something makes me think that 'GetRight' tool might have been able to do the same thing, but I can't seem to see the feature on their website.
- flatroze 7y agoAh, I remember using something like that. I thought that tool was saving it into one .html file, but data URLs didn't exist back then, so creating directories alongside with HTML files was the only option to "replicate" a web resource, now I understand exactly what you were talking about. I'll do some more digging around and implement that in the nearest future. I may need to make all the requests async first to make sure that saving one resource with decent depth won't take too long.
- olakeasseillo 7y agoyes, "teleport pro" for win98, you could scrape a site and duplicate locally or scrape only for specific file type or size, had recursive link follow depth option and created several threads for the requests(sniff)
- leshokunin 7y agoThis would be a perfect fit for IPFS. I love the idea of having just one file in a permanent link.
- flatroze 7y agoThis could also be an interesting alternative to PDF, especially with web fonts embedded as data URLs.
- turbinerneiter 7y agoI've been using this kind of standalone PDFs produced from Markdown with pandoc for a while, and the possibilities are insane. Imagine a paper in the form of a single HTML file, which has (a subset of) the data included, the graphs zoomable, the colors chanegable (to whatever vision problems you have) - maybe even the algorithm to play around with! Jupyter Notebooks already go in that way. only without the single-file, open in browser aspect, I think.
- dfee 7y agoIf the output were a tar file, couldn’t we also say it was saving web pages as a single file? Wouldn’t that also be easier?
- flatroze 7y agoI think there's an issue with opening a tar file, e.g. if sent to someone who needs to view the document but isn't techy. It seems to me that having one file that any browser can easily open (and not require Internet connection to view) is a big advantage over having a directory with assets alongside the .html file. It may be one of those things that make things easier yet nobody really complains about how things are usually done when the page gets saved. I hope more browsers add support for saving pages as MHTML in the nearest future so that we wouldn't need tools like this one.
- Springtime 7y agoMHTML is pretty good for this already btw (not to take away from this neat project though :)). Similarly stores assets as base64'd data URIs and saves it as a single file. Can be enabled in Blink-based browsers using a settings flag and previously in Firefox using addons (also in the past natively in Opera and IE).
- flatroze 7y agoApparently everybody knew about MHTML but me Ü I'm going to look into that format and see if I could enhance monolith to output proper MHTML, among other additions and improvements. Thank you for the info!
- masklinn 7y agoI don't know that it would be a very useful thing to do at least in the short term: there's a bunch of "web archive" formats out there and the common thread between them is that they're custom archive formats, you need special clients or support for those formats: * mthml encodes the page as a multipart MIME message (using multipart/related), essentially an email (you're usually able to open them by replacing the .mth by .eml) * WARC is its own thing with its own spec * WAFF is a zipfile, not sure about the specifics * webarchive is a binary plist, not sure about the specifics either Your tool generates straight HTML which any browser should be able to open. It probably has more limitations, but it doesn't require dedicated client / viewer support. Maybe once you've got all the fetching and extracting and linking nailed down it would be a nice extension to add "output filters", but that seems more like a secondary long-term goal, especially as those archive formats are usually semi-proprietary and get dropped as fast as they get created (WARC might be the most long-lived as it descends from the Internet Archive's ARC, is an ISO standard and is recognised as a proper archival format by various national libraries).
- mftrhu 7y agoThere isn't much to WAFF. Each WAFF file can contain more than one saved page. Each page needs to be contained within its own folder (whose name is usually the timestamp of when the page was saved, but it doesn't matter AFAICT). There can be an `index.rdf` file in there, to specify metadata and which file to open, but otherwise you should look for an `index.SOMETHING` file - usually `index.html`. E.g. test.maff `-- 1566561512/ |-- index.rdf |-- index.html `-- index_files/ `-- ??? When I was messing around with archiving things locally I settled on WAFF, because it's pretty much trivial to create and to use. Even if your browser does not support it, you just need to unpack it to a tempdir and open the index file.
- alpb 7y agoI think it would be way better to explain in the repository: - how do you handle images? - does it handle embedded videos? - does it handle JS? to what extent? - does it handle lazily loaded assets (i.e. images that load only when you scroll down, or JS that loads 3 seconds later after the page is loaded) In general, how does this work? The current readme doesn't do a decent job explaining what the tool exactly is. For all I can tell, it probably just takes a screenshot of the page, encodes as base64 into the html and shows it.
- flatroze 7y agoGood points, thank you for the review. I'll work on enhancing the readme file to be more informative.
- deleted 7y ago[deleted]
- quickthrower2 7y agoIt can’t handle JS completely because we can’t predict a programs behaviour using static analysis. See Halting Problem for example.
- kuzehanka 7y agoI saw a tool that handles JS to a limited extent by capturing and replaying network requests to accommodate said JS. It records your session while you interact with a site, and is then able to replay everything it captured. This tool was able to capture three.js applications and other interactive sites quite well.
- VvR-Ox 7y agoVery cool idea - thank you for this! On question: How does it handle those cookie pop-ups, gdpr-warnings etc?
- flatroze 7y agoOh, thank you kindly. That's an interesting question. I think it depends on how the given modal is implemented, but closing them should technically work (unless the page is saved with JavaScript code removed [-j flag]). Those notifications can easily be removed from the saved file using any text editor, should be pretty easy if you know how to edit HTML code. I don't think removing it would violate anything since "this website" will no longer really be a website but rather a local document at that point.
- Exuma 7y agoSaving this for later
- dredmorbius 7y agoFYI: "favorite" is one way of doing that through HN. Bookmarks, or downloads, externally.
- skinnymuch 7y agoFavorites is limited to a certain amount on HN before you start losing the oldest favorite.
- dredmorbius 7y agoHow many specifically? I'm at 252 posts, presently, just checked. That seems to be a complete log.
- sah2ed 7y agoIs this limitation documented somewhere? After how many entries before the HN software started tripping on your favorites?
- dang 7y agoThat's not true. One user has 46,000. What did you see that made you think this?
- dvcrn 7y agoI'm not so experienced but how does this compare to .webarchive?
- flatroze 7y agoThe idea is almost identical, yet saving as .webarchive is only supported by Safari, and it's also not a plaintext format, hence can't be edited as easily.
- deleted 7y ago[deleted]
- sbmthakur 7y agoNice work! I am wondering if Puppeteer can also be used to accomplish the same thing.
- flatroze 7y agoIt for sure would help with those SPA websites that get their DOM fully generated by JS. A web extension that saves the current DOM tree as HTML would perhaps do a better job, especially when it comes to resources which require some web-based authentication.
- fouc 7y agoSweet idea! I would especially like to be able to capture videos and pictures too. I suspect for saving videos, a good approach would be some sort of proxy + headless browser combination, where the proxy is responsible for saving a copy of all data the browser requests for. Thoughts?
- flatroze 7y agoThanks! Pictures should work, I'll check more tags first thing tomorrow when I start working on improving it. I use youtube-dl for youtube and other popular web services myself. Embedding a video source as a data URL could in theory work, but it'd be quite a long base64 line. Also, editing .html files with tens or hundreds of megabytes of base64 in them would perhaps be less than convenient.
- makach 7y agoAhh, to me it looks like it creates an amalgamation of the web page+contents. How does this work on neverending webpages/forever scroll? How will it behave if you need to authenticate before browsing the page?
- flatroze 7y agoThat's it in the nutshell! It seems to work for basic pages quite well, I think that lazy load will work for most pages as long as the JavaScript is embedded (no -j flag provided) and the Internet connection is on. It saves what's there when the page is loaded, the rest is a gamble since every website implements infinite scroll differently. Authentication is another tricky part -- it's different for every browser. I will try to convert it into a web extension of sorts, so that pages could be saved directly from the browser while the user is authenticated.
- donatzsky 7y agoFor authentication, you could add an option for passing http headers, as well as accept Netscape-style cookie files. Whenever I want to download a video, using YouTube-dl, from a site that requires authentication, I first login using my browser and then exports the cookies using an extension.
- sah2ed 7y agoMay I ask what extension you use for cookie exporting?
- tenken 7y agoHow is this different from https://en.m.wikipedia.org/wiki/Web_ARChive https://en.m.wikipedia.org/wiki/Web_ARChive
- masklinn 7y agoIt looks like it creates a normal HTML file (embedding assets as data URI) so it should require no special client / support. HTMLD, WARC, MTHML, MAFF and webarchive are all "container" formats which bundle assets next to the HTML using various methods (resp. bundle, custom, multipart MIME, zip and binary plist).
- emerongi 7y agoThe issue with this is that if the website requires some external API for content, it might not work properly. https://webrecorder.io/ https://webrecorder.io/ solves that problem by recording all interactions and then replaying them as needed. > Webrecorder takes a new approach to web archiving by “recording” network traffic and processes within the browser while the user interacts with a web page. Unlike conventional crawl-based web archiving methods, this allows even intricate websites, such as those with embedded media, complex Javascript, user-specific content and interactions, and other dynamic elements, to be captured and faithfully restaged.
- deleted 7y ago[deleted]
- jplayer01 7y agoAh, I've been thinking about making something like this. You beat me to it. I've been using the SingleFile add-on until now. I'll definitely give this a try.
- jordwalke 7y agoI really like this concept, and I've been using an npm package called inliner which does this too: https://www.npmjs.com/package/inliner https://www.npmjs.com/package/inliner I'm glad there's more people taking a look at the use case, and I'd be interested to see a list of similar solutions. If you combine this with Chrome's headless mode, you can prerender many pages that use JavaScript to perform the initial render, and then once you're done send it to one of these tools that inlines all the resources as data URLs. /Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome ./site/index.md.html --headless --dump-dom --virtual-time-budget=400 The result is that you get pages that load very fast and are a single HTML file with all resources embedded. Allowing the page to prerender before inlining will also allow you to more easily strip all the JavaScript in many cases for pages that aren't highly interactive once rendered.
- flatroze 7y agoThanks, I'll look into making it work with pipes or some other way to interact with headless browsers.
- ahub 7y agoReach out when you do so, I've a similar use case here !
- gildas 7y agoYou want this: https://github.com/gildas-lormeau/SingleFile/tree/master/cli https://github.com/gildas-lormeau/SingleFile/tree/master/cli.
- sergioisidoro 7y agoThis sounds great, but the first thing I thought was how this would be a perfect tool to make automated mass phishing scams. If the outcomes are realistic, take a massive list of sites, make a snapshot of each page, replace the POST login URLs with the phishers, deploy these individual HTML files, and spread the links through email. I wonder how does this project handle forms.
- flatroze 7y agoThank you for reminding me, I need to set action="" to be an absolute path when the page is saved. upd: Done, now forms get their action="/submit" converted into action="https://website.com/submit" https://website.com/submit" when the page is saved.
- danfang 7y agoI'd imagine you can already do that with some basic web scraping tools. This definitely makes it easier.
- js8 7y agoI am using "Save Page WE" Firefox extension for this. Better at saving JS content and less clutter than saving all the images and stuff.
- personjerry 7y ago`cargo install` install 237 packages for this?! I don't think that's acceptable.
- Deukhoofd 7y agoProbably the reqwest crate. That thing alone uses like 30 crates, not including the dependencies of those crates.
- flatroze 7y agoThe compile time is rather long as well, I'm looking into ways of reducing the amount of dependencies.
- codegladiator 7y agoThen don't accept it. What's with the number of packages or size ?
- mikekchar 7y agoWith respect to the Unlicense, does anybody have any knowledge about how good it is in countries which don't allow you to intentionally pass things into the public domain (most countries that aren't the US)? How does it compare to CC0 in that respect?
- flatroze 7y agoWhich license would you recommend to release this software under to reach the broad adoption yet permissive terms, if not the Unlicense?
- mikekchar 7y agoI honestly don't know. That's why the question :-) Is CC0 good for software? It seems to be a bit more complete from a non-US view point, but I don't know if there are lurking situations. Possibly MIT is better -- it's pretty darn permissive. I'm really just soliciting opinions.
- pgcj_poster 7y agoYes, CC0 is the only Creative Commons License suitable for software. It's endorsed by the Free Software Foundation [1], although not the Open Source Initiative. I use it for everything that I don't want to copyleft. [1] https://www.gnu.org/licenses/license-list.html#PublicDomain https://www.gnu.org/licenses/license-list.html#PublicDomain
- hendry 7y agoI imagined that https://www.w3.org/TR/widgets/ https://www.w3.org/TR/widgets/ would be the open container format for saving a Web app to a single file.
- personjerry 7y agoStyling breaks on this site: https://www.scientificamerican.com/article/the-hunt-is-on-for-alpha-centauris-planets/ https://www.scientificamerican.com/article/the-hunt-is-on-fo...
- flatroze 7y agoThank you for the heads up, I'll test it and enhance to preserve styles better.
- FreeHugs 7y agoOne thing I always wonder when I see native software posted here: How do you guys handle the security aspect of executing stuff like this on your machines? Skimming the repo it has about a thousand lines of code and a bunch of dependencies with hundreds of sub-dependencies. Do you read all that code and evaluate the reputation of all dependencies? Do you execute it in a sandboxed environment? Do you just hope for the best like in the good old times of the C64?
- jasonvorhe 7y agoWhat computer are you using and which operating system is running on that? Have you read the code?
- interfixus 7y agoWe may semi-trust our package and repo systems. This tool is readily available through AUR on my Arch machine, I see. Or we may go the whole hog and actually have a peek through the source.
- laumars 7y agoAUR are packages that aren't in Arch's repo system. Granted tools like yaourt do make installing AUR packages nearly as easy as pacman but anyone can upload anything to AUR thus you are expected to vet the packages yourself (hence why tools like yaourt repeatedly prompt you to read the build scripts et al before running them).
- ahub 7y agoI noticed there is a `-j` argument to remove javascript. A `-i` argument for removing images would be great too.
- flatroze 7y agoIt is done, option -i in the latest version (2.0.3) now replaces all src="..." attributes with src="<data URL for a transparent PNG pixel>" within IMG tags.
- interfixus 7y agoNice. I can see some automated uses for this. In ordinary browsing, am currently using a Firefox addon called SingleFile which works surprisingly well. Stuffs everything into (surprise, surprise) one huge single file - html with embedded data, so compatible everywhere.
- flatroze 7y agoIt sounds like a great add-on, I have to check it out to see what it does to remote assets and how it works with asynchronously loaded assets.
- tannhaeuser 7y agoWell you could do that for a long time with MHTML, WARC, etc. downloaders, including those available in browsers via "Save Page as", though CSS imports aren't covered by older tools (are they by yours?). Anyway, congrats for completing this as a Rust first-timer project, which certainly speaks to the quality of the Rust ecosystem. For using this approach as offline browser, of course, the problem is that Ajax-heavy pages using Javascript for loading content won't work, including every React and Vue sites created in the last five years (but you could make the point those aren't worth your attention as a reader anyway).
- flatroze 7y agoCSS imports are covered by converting .css files into data URLs, later I will parse those and embed resources found within stylesheets as well.
- lucasverra 7y agosuper project ! i ve pretty baffled with the difficulty to save a webpage in proper format. I’ve tried with PDF converter, getPolaroid app and of course firefox screen shot feature for the entire scroll thing. Will try this for saving purposes. I am also interested in cloning/forking sites for modification purposes, I will feedback you on the results four my consulting gigs
- flatroze 7y agoThank you for the kind words. It will evolve into a reliable tool in a couple weeks and it should eventually work for embedding everything, including things like web fonts and @url()'s within CSS. If anything doesn't work, please open an issue, I have plenty of time to work on it.
- fit2rule 7y agoI've been printing to PDF for decades now, and nothing comes close to the ease of use and versatility of 2 decades worth of interesting web pages .. I have pretty much every interesting article, including many from HN, from decades of this habit. Need to find all articles relating to 'widget'? $ ls -l ~/PDFArchive/ | grep -i widget This has proven so valuable, time and again .. there is a great joy in not having to maintain bookmarks, and in being able to copy the whole directory content to other machines for processing/reference .. and then there's the whole pdf->text situation, which has its thorns truly (some website content is buried in masses of ad-noise), but also has huge advantage - there's a lot of data to be mined from 50,000 PDF files .. Therefore, I'd quite like to know, what does monolith have to offer over this method? I can imagine that its useful to have all the scripting content packaged up and bundled into a single .html file - but does it still work/run? (This can be either a pro or a con in my opinion..)
- flatroze 7y agoI'd say since monolith produces a plaintext document it lets you edit things easier if needed. JS can be removed from the final document using the -j flag. HTML Files can also be grepped for content, unlike PDFs.
- dredmorbius 7y agoHaving gone this route in part myself, advantages of HTML or other more-structured file formats, if there is appropriate metadata markup: - Allow for recording source and author information (PDF ... doesn't always provide this). - Allows for full-text search. - Allows for editing out annoyances. I'll frequently go from HTML to some simplified representation (e.g., Markdown), and then re-generate formats that are useful elsewhere: HTML, PDF, ePub, etc. Dumping from HTML to Markdown frequently makes cruft-removal far simpler, and the principle content of most pages is text. In rare instances, images are useful, and even more rarely, any multimedia content (video, audio, programmatic content). What's depressing is the number of sites which screw with even basic HTML. E.g., the NY Times rarely use HTML tables for tabular representation, and instead use a homebrew combination of custom markup, CSS, and JS to much the same effect. Pretty, in situ, but brittle and transports exceedingly poorly. That's just one of many such cases.
- gildas 7y agoNote that SingleFile can easily run on command line too, cf. https://github.com/gildas-lormeau/SingleFile/tree/master/cli https://github.com/gildas-lormeau/SingleFile/tree/master/cli.
- mikaelmorvan 7y agoThe main problem with your code is that you only handle simple web1 site. What about javascript execution ? If you replay your capture, you have no idea of what you will see on general Web2 website. The only way I know to capture a web page properly is to "execute" it on a browser. Gildas, the guy behind SingleFile (https://github.com/gildas-lormeau/SingleFile https://github.com/gildas-lormeau/SingleFile) is well aware of that and his approach realy works everytime. Try on a Facebook post, a Tweet, ... It just works.
- deleted 7y ago[deleted]
- lucideer 7y agoThe capture includes JS, so this should work for most JS-dependent sites, with the exception of scripts loading other additional assets. Tbh, often those are superfluous, or egregious examples of bad web dev, so it seems a reasonable solution for most cases. SingleFile is a different approach, but it's a lot more involved/less convenient than a cli, and loading in something like WebDriver on the cli for this would be overkill, unless you're doing very serious archival work.
- mikaelmorvan 7y agosuperfluous, or egregious examples of bad web dev?? Do you know what Web 2.0 is? Do you know what are React, Angular, and the other JS Frameworks? When you create a modern webapp, a lot of data are retrieved from servers as Json and formated in the browser in Javascript. Even sometimes Css is generated on browser-side. Even more, on webapp where user login is taken into account, the display is modified accordingly. That's the web of 2019. The approach consisting of geting remote files and launching them in a browser is really naive. Speaking of SingleFile, it as a cli version and can handle full web 2.0 webapp without any problem. And of course, the Web 1.0 webapps work as well.
- tinsx 7y agoI think that's exactly what that person means by superfluous and egregious examples of bad web development; SPAs, javascript frameworks of that nature. :p
- nessunodoro 7y agocall me old fashioned, but I still use Ctrl+S
- ur-whale 7y agoDoes not compile with some byzantine message about let in const funcs being unstable.
- flatroze 7y agoCould you please open issue on github providing the output that you get in the terminal?
- ur-whale 7y agoI have closed my github account since the takeover occured.
- dspillett 7y agoThe important part of what he said was "providing the output that you get in the terminal". Simply stating "I got an error" and expecting the developer(s) to use clairvoyance to glean further detail is far from a helpful way to report a problem. Perhaps dropping the details in a pastebin site and linking to that would be a possible alternative? Or just including the error message here if it is short enough, though HN shouldn't really be used as a tech support channel.
- mrieck 7y agoIf you only want a portion of a webpage I made a tool called SnipCSS for that: https://www.snipcss.com https://www.snipcss.com The desktop version saves an HTML file, stylesheet and images/fonts locally, and it only contains the HTML of the snippet with the CSS rules that apply to the DOM subtree of the element you select. I'm still working out bugs but it would be great if people try it out and let me know how it goes.
- sansnomme 7y agoThere have been quite a few extensions in this space: https://stackoverflow.com/questions/10266334/add-on-to-copy-a-page-element-with-styles https://stackoverflow.com/questions/10266334/add-on-to-copy-... https://github.com/Dalimil/Web-Design-Pirate https://github.com/Dalimil/Web-Design-Pirate
- mrieck 7y agoI tried SnappySnippet before when looking into the idea - it didn't work well for me and crashed often. I never saw DesignPirate, but just now I tried it and it didn't output any CSS. I'm not sure but it doesn't look like either of these use chrome.debugger API to call devtools api methods. (you get a warning in Chrome if you use that) I'm hoping my tool will be better so it's good enough people would be willing to pay for it, but we'll just have to see.
- dtjohnnymonkey 7y agoThank you for this. I’ve been looking for something that does this exact thing. I don’t like any of the other HTML archiving formats .
- sametmax 7y agoGood, but won't work with the heavy JS pages using Ajax to load any single content. The firefox extension seems to do that : https://addons.mozilla.org/fr/firefox/addon/single-file/ https://addons.mozilla.org/fr/firefox/addon/single-file/
- gildas 7y agoUnfortunately it's not written in Rust so it won't make the first page of HN.
- cr0sh 7y agoThis is interesting - I think any of us who save things off the internet have made something like this (I usually save entire sites or large chunks, though - so I have a different toolset - still, I also do single pages, so I might try out this tool). One thing I would propose to add - either a flag, or by default - have it parse the path to the page and create the file with the name - that way you can just "monolith {url}" and not have to worry about it. I am also curious as to how it handles advertisements and google tracking and such; some way to strip out just those scripts (and elements) could be handy.
- sankalp210691 7y agoThis is pretty useful. It would be great to have a functionality of converting the HTML page to a PDF as well.