15 ms·
Ask HN: What is nowadays (opensource) way of converting HTML to PDF?
I'm using wkhtmltopdf but it is painful to work with? what are other people using nowadays? i.e canva or other tools?
- zja 1y agopandoc
- hhthrowaway1230 1y agodoesn't pandoc rely on some engine itself?
- brudgers 1y agoCurious why that matters to you? I mean everything has dependencies (some of the solutions elsewhere require Chrome and other common solutions require the JVM). At least Pandoc is GPL.
- kakokiyrvoooo 1y agoIt matters because pandoc is not rendering the website to pdf, it converts the html to latex and then uses a latex engine to render the pdf.
- brudgers 1y agoForgive me but I don’t understand why that matters to you and am trying to understand what the issue with Latex is. Because lots of things work this way. For example compilers built on LLV uses an intermediate language and Python uses byte code. I suspect some html to pdf tools go through postScript.
- kreetx 1y agoThere are multiple ways to "depend", so if pandoc executes some external tool all of the work then might as well use that external tool directly. You will get more control over how the conversion happens, know for what search for when in trouble etc.
- brudgers 1y agoMy understanding and experience is that Latex has a significant learning curve and Pandoc provides a more gentle front end. Of course Latex gives you fine control to hand tune the engine…but that doesn’t seem like what the OP is looking for.
- kreetx 1y agoSure, I don't mean that anyone would look at the Latex in between. I'm just saying that if tool x directly calls tool y to do the job then might as well use tool y directly.
- brudgers 1y agoSince hammers and nails are a common tool-workpiece example…consider the nail gun. Theoretically you can drive nails with a 22 caliber blank cartridge without making the “call” through a nail gun. But you won’t finish laying shingles as quickly and easily… Or to put it another way, there’s a reason assemblers are almost always better than machine code and compilers are almost always better than assemblers for the ends people care about. I mean why use Latex at all when you could write your own typesetting language? Maybe because you are not a knuth.
- kreetx 1y agoYou're confusing wrappers with alternatives. The comparison is more like if somebody published a script called html-to-pdf.sh which directly calls, e.g, chrome, would you want to use this script or use chrome directly? I would prefer the latter because (1) I would know what actually does the conversion, (2) I would know what to search for on the web should I need to tweak the output. This knowledge gives me more power as I know the actual converter. The wrapper script perhaps only helps with what the command line should be.
- cpach 1y agoYep, you need something like XeTeX in order to render the PDF.
- beeforpork 1y agoDoes pandoc do JavaScript? For stuff that is rendered (I don't want animated, interactive PDFs...).
- w10-1 1y agoTo reinforce this: pandoc has been the go-to for a long, long time and they have encountered and addressed tons of issues, which is especially important for two underspecified and over-provisioned formats like HTML and pdf. Go through the revision and bug history to see a sample of issues you're avoiding by using a highly-trafficked, well-supported solution. The only reason not to use it is when they say they don't support a given feature that you need; and the nice thing there is that they'll usually say it, and have a good reason why. The other reason to use pandoc is that while you might currently want PDF as your outbound format, you might end up preferring some other format (structured logically instead of by layout); with pandoc that change would be easy. Finally, pandoc is extensible. If you do find that you want different output in some respect, you can easily write an plugin (in python or haskel or ...) to make exactly the tweak you need.
- throw03172019 1y agoI run chromium on my server and render the PDF from there using puppeteer.
- kappadi3 1y agoPuppeteer and Playwright are the main open-source options nowadays, both solid for HTML → PDF once your print CSS is sorted. Don’t forget proper page breaks (break-before/after/inside) — e.g. break-after: page works in Chromium, while always doesn’t. For trickier pagination you can look at Paged.js, and I’d test layouts in Chrome/Edge before automating. Shameless plug: I run yakpdf.com, a hosted Puppeteer-based service if you want to avoid self-hosting. https://rapidapi.com/yakpdf-yakpdf/api/yakpdf https://rapidapi.com/yakpdf-yakpdf/api/yakpdf
- johnh-hn 1y agoSeconded. I went with C# + Playwright. I tried iTextSharp, iText, PDFSharp, and wkhtmltopdf, but they all had limitations. I had good results with Playwright in minutes, outside of tweaking the CSS like you mention. I documented the process here[0] if anyone needs examples of the CSS and loading web fonts. Apologies for the article being long-winded – it was the first one I published. [0] https://johnh.co/blog/creating-pdfs-from-html-using-csharp https://johnh.co/blog/creating-pdfs-from-html-using-csharp
- benoau 1y agoThirded, you can build this straight into your backend or into a microservice very easily. You can also easily generate screenshots if that's more suitable than PDFs. You can also easily use this to do stuff like jam a set of images into a HTML table and PDF or screenshot them in that format.
- ChuckMcM 1y agoYou made me realize that tractor feed roll paper would be really great for printed web pages, no page breaks! Kinda like reading scrolls of yore.
- mightjustwork 1y agohttps://gotenberg.dev/ https://gotenberg.dev/ ...has been working well for me for the last few years. It's a headless instance of Google Chrome with a golang wrapper. Runs well in Docker or a cloud instance.
- pabs3 1y agoJust print to PDF in a browser, or automate that using a browser automation tool. For a non-browser-based open source solution, WeasyPrint. https://weasyprint.org/ https://weasyprint.org/ For a proprietary solution, try Prince XML: https://www.princexml.com/ https://www.princexml.com/
- rossdavidh 1y ago+1 to weasyprint; I have used weasyprint with a django production system for a few years now, and it works well enough that I never have to think about it. I'm not doing anything fancy, though, but for me it has worked well.
- grounder 1y agoWeasyPrint works really well for me. It can support all of the languages and fonts I need. I run it on AWS Lambda and in Docker as a web service. I previously used WKHTMLTOPDF, but it hasn't been supported for years and doesn't support the latest CSS, etc. It does support JS if you need it, but I'd probably look at headless Chromium or another solution for JS if needed. Edit: Previous post with some good discussion: https://news.ycombinator.com/item?id=26578826 https://news.ycombinator.com/item?id=26578826
- stuaxo 1y agoThis is my experience and recommendation too.
- sureglymop 1y agoPrince XML looks nice but what about creating a PDF directly from a website? This often adds some problems, for example links still pointing to other pages on the web. But in my experience printing to PDF is often not good enough.
- chinathrow 1y agoYes, I did that for a recent small program. The @media print media query is powerful enough for most of the stuff I wanted to format nicely. Even page breaks are possible.
- journal 1y agoif you are doing html to pdf, you might also need the ability to merge. a few more features and you're better of with a commercial solution.
- crazygringo 1y agoMerge what?
- pentium166 1y agoI assume combining 2+ documents. For example, attaching a cover page with document owner/version control/lifecycle information to an existing PDF.
- crazygringo 1y agoThat's the easiest thing in the world with free software. One way is to install poppler-utils and use pdfunite. There are many other open-source packages you can use as well.
- fogzen 1y agoDon’t. Show a web page and open the print dialog, and tell people to save as PDF. All major browsers support this, and the browser HTML to PDF code is the most robust and accurate.
- crazygringo 1y agoThere's nothing in OP's question that suggests this is a one-off operation in response to a user action. It's very likely to be a massive batch operation of a ton of HTML files that might not even be their own site.
- hhthrowaway1230 1y agothis is the case indeed
- chibbell 1y agoThat does make sense where possible. I do feel like OPs question is super relevant if you are doing anything where the PDF has to be rendered server side, like say as part of a larger data process when producing an exportable report in PDF format.
- Snawoot 1y agochrome --headless --disable-gpu --print-to-pdf https://example.com https://example.com
- piptastic 1y agosame: google-chrome --headless --disable-gpu --no-pdf-header-footer --hide-scrollbars --print-to-pdf-margins="0,0,0,0" --print-to-pdf --window-size=1280,720 https://example.com https://example.com ended up using headless chrome specifically to make sure javascript things rendered properly
- hhthrowaway1230 1y agoUsed this, sigh of relief, thank you
- HPsquared 1y agoCan Chromium do this? Edit: it appears so- https://news.ycombinator.com/item?id=15131840 https://news.ycombinator.com/item?id=15131840
- nine_k 1y agoYes, routinely works for me.
- mmphosis 1y agoCan Firefox do this? with an elaborate script that relies on xdotool
- andrehacker 1y agoYes, kind of... /path/to/firefox --window-size 1700 --headless -screenshot myfile.png file://myfile.html Easy, right ? Used this for many years... but beware: - caveat 1: this is (or was) a more or less undocumented function and a few years ago it just disappeared only to come back in a later release. - caveat 2: even though you can convert local files it does require internet access as any references to icons, style sheets, fonts and tracker pixels cause Firefox to attempt to retrieve them without any (sensible) timeout. So, running this on a server without internet access will make the process hang forever.
- exabrial 1y agoopenhtmltopdf is what we're using. Some outdated versions.
- supersaw 1y agoBeen using this as well. It's worth noting that while the original project appears to have been abandoned, it has since been forked and is currently maintained here: https://github.com/openhtmltopdf/openhtmltopdf https://github.com/openhtmltopdf/openhtmltopdf
- exabrial 1y agothanks, didnt know that!
- RiverCrochet 1y agoIf you don't really need the PDF but just want to archive pages, SingleFile is better. It'll capture the entire page to a single HTML file and I find this is better than the PDF if I don't want to print it. It's a browser extension, but there's also a command line version (https://github.com/gildas-lormeau/single-file-cli https://github.com/gildas-lormeau/single-file-cli) that uses Chrome or Chromium's headless mode.
- haft 1y agoA revers of this question; what is the best way to convert pdf to html? We are required by accessibility law to make our PDFs WCAG compliant however it would be easier to convert these to HTML.
- haft 1y agoA reverse of this question; what is the best way to convert pdf to html? We are required by accessibility law to make our PDFs WCAG compliant however it would be easier to convert these to HTML.
- bencornia 1y agoI have been using pdf2htmlex with some success. https://github.com/pdf2htmlEX/pdf2htmlEX https://github.com/pdf2htmlEX/pdf2htmlEX
- drabbiticus 1y agoThis is really cool, so thanks for sharing. Since the motivating goal for the question you are answering is WCAG compliance, is the output of pdf2htmlex meaningfully more WCAG compliant?
- fschuett 1y agoRendering to SVG, at least that's what I did on https://fschutt.github.io/printpdf/ https://fschutt.github.io/printpdf/ I am currently writing a WASM-ready PDF toolkit that can handle both HTML to PDF and then rendering PDF pages to SVG. However, it's not yet production-ready. The underlying HTML engine is currently a severe "work in progress", but it gives me the low-level access that I need: https://azul.rs/reftest https://azul.rs/reftest
- dredmorbius 1y agoThis reduces to parsing PDFs, which is an unsolved hard problem. At low volumes, my preferred approach is to select and extract text (copy/paste, perhaps using the poppler library for larger-scale work), dump that to plain-text and convert that (manually / scripted) to Markdown. From there you can get to PDF or pretty much any other format through tools such as pandoc.
- etyhhgfff 1y ago[dead]
- pentium166 1y agowkhtmltopdf is pretty out of date at this point and headless Chrome/Chromium or something that wraps them is probably a better and safer, roughly equivalent, alternative. Docker might not be a great option if you're already running a containerized service and don't want to deal with getting them to play nice together.
- Aachen 1y agoPlease don't turn nice formats into a format that's similar to screenshots of text. Pandoc has an option to pack all images and styles needed to render the page into one html file: pandoc --self-contained input.html -o output.html
- agedclock 1y agoPandoc would be my preferred tool. It is excellent at converting between other formats as well.
- TylerE 1y agoBeing (not so easily) edited is often a feature, not a bug.
- guywithahat 1y agoI was thinking this too, PDF's exist so people don't mess with the document. That said, it's still a clever feature, and pandoc can convert html into a pdf as well with a conversion engine. That said, I suspect it'll fail on anything sufficiently complex pandoc input.html -o output.pdf --pdf-engine=<your engine>
- ryandrake 1y agoIs this really that much of a motivation in 2025? Maybe in 2000 you could publish a PDF with the assurance that only the people who paid for Acrobat would be able to edit it, but today, there are a lot of accessible ways to edit PDFs, I don't think I'd choose PDF if I for whatever reason wanted to limit others from editing.
- craftkiller 1y agoIf that is your goal, you should be cryptographically signing your documents with your PGP key. That way you actually have assurance the document has not been modified rather than just hoping someone hasn't modified the document. Additionally, PGP can sign anything so you are open to use whatever format you want.
- Aachen 1y ago
- thangalin 1y agoIs this an xy problem? If you have the original document (in Markdown), one possibility would be to use my software, KeenWrite[1], to convert Markdown to XHTML then typeset XHTML to PDF via ConTeXt. See the user manual[2] for an example of a Markdown document typeset in this fashion, along with usage instructions. If you only have HTML to work with, you can also use Flying Saucer[3], which is what KeenWrite uses to preview Markdown documents when rendered as HTML. Flying Saucer uses an open-source version of iText[4] to produce PDF documents (from HTML source docs). Another possibility is to use pandoc and LaTeX. [1]: https://keenwrite.com/ https://keenwrite.com/ [2]: https://keenwrite.com/docs/user-manual.pdf https://keenwrite.com/docs/user-manual.pdf [3]: https://github.com/flyingsaucerproject/flyingsaucer https://github.com/flyingsaucerproject/flyingsaucer [4]: https://itextpdf.com/ https://itextpdf.com/
- nicoburns 1y agohttps://github.com/plutoprint/plutobook https://github.com/plutoprint/plutobook was a recent Show HN and looks excellent
- deleted 1y ago[deleted]
- ftchd 1y agothe only thing I found to work reliably well is simply Chromium's print feature
- hhthrowaway1230 1y ago5k pdfs a month for archival purposes, must be pdf, customers demand this
- bob1029 1y agoIf your HTML is simply an intermediary to get you to a PDF, you could consider just skipping straight to building the PDF directly: https://pdfbox.apache.org https://pdfbox.apache.org This would be far more efficient than spinning up an entire browser and printing PDFs to disk.
- deaddodo 1y agoBuilding PDF directly (unless you're creating documents, especially fillables) is non-intuitive. Most PDFs are people trying to capture live data in a cached manner. If not, using a preliminary format like Markdown/HTML/LaTeX/DocX/etc to generate your PDF is almost always more intuitive.
- ratStallion 1y agoMy website's content is xml, and I use Apache Fop to turn it into a PDF with page numbers and other nice things. It works nicely, but takes some setup.
- juice_bus 1y agoI have Chromium shoved into an AWS Lambda Layer, when we need HTML to PDF conversion we shove it off onto that. It loads the HTML into Chromium then "prints" it to PDF.
- freedomben 1y agoI'd love to go the other way: convert a PDF into a self contained HTML page that renders properly in a browser. It's been way harder than I thought it would. Any advice?
- drabbiticus 1y ago> renders properly Depending on your requirements on both PDF input and HTML output, there is often no way to do this that is both easy and general. At it's core, PDFs are not designed to be universally reflowable.
- mr_mitm 1y agoYou could embed it as a base64 blob, embed PDF.js (which is included by browsers anyway, I think) and use that to render it in the HTML. But I realize you probably meant a static HTML without JavaScript.
- freedomben 1y agoYes ideally, but even this is helpful, thank you!
- gucci-on-fleek 1y agoYou can use dvisvgm, pdftocairo, or Inkscape to convert PDF to SVG, which you can either use directly or insert inline into an HTML document.
- lizimo 1y agoIf generating PDF dynamically is what you really care about, consider Typst. https://typst.app/ https://typst.app/ We use it in production to generate reports, and it is amazing.
- leephillips 1y agoSee https://lwn.net/Articles/1037577/ https://lwn.net/Articles/1037577/ for a recent summary of what you can do with Typst.
- lovelydata 1y ago[dead]
- mococa 1y agohttps://gotenberg.dev https://gotenberg.dev
- roxolotl 1y agoWould second this. I’ve used it in production to generate tens of thousands of PDFs a day. It just works. Run the docker container throw html and variables at it and get PDFs back.
- NanoWar 1y agoWe use this in production and it's very stable. It also supports background gradients which we wanted to use so badly :-) Can recommend
- Glyptodon 1y agoThe last time I had to do this I scripted a back-end that scaled up headless chrome browsers to render web pages to PDF. I think it was using Puppeteer, but was a few years ago. (FWIW the decision I think was mostly driven by the environment, I think there are other options.)
- gigatexal 1y agopandoc is your friend.
- lucis 1y agojsPDF is a work of art https://parall.ax/products/jspdf https://parall.ax/products/jspdf
- flanbiscuit 1y agoBeen looking at this one. I inherited a project and I set it up to use puppeteer and chrome server side to generate a PDF from HTML but it's too much overhead. I want to do this all on the frontend because it should be simple enough to do and can use less resources on the server.
- handzhiev 1y agoI'm surprised no one mentioned mPDF. Maybe php isn't very popular here :)
- cjm42 1y agoI've had decent results with html-pdf-chrome[0], which automates printing to PDF from Chromium or Chrome. [0] https://github.com/westy92/html-pdf-chrome/ https://github.com/westy92/html-pdf-chrome/
- syngrog66 1y agopandoc
- efnx 1y agopandoc
- ineedasername 1y agoGhostscript. Depending on specific needs it may be much more turnkey than Pandoc, which isn’t actually doing much directly with things other than intermediating, iiuc. (LaTex) does the heavy lifting. Ghost script is working with postscript natively and will likely manage idiosyncrasies of web content better. It’s got a decent ecosystem, command line, you can find gui’s if that’s your thing (no judgement, your lifestyle is none of my business). Many other good tools mentioned here as well, but if your asking because you need more, or fine grained (near infinite) control over the pdf composition, there’s nothing OSS I can think of that approaches its capabilities. https://ghostscript.com/ https://ghostscript.com/
- roschdal 1y agoOpenPDF for Java https://github.com/LibrePDF/OpenPDF https://github.com/LibrePDF/OpenPDF
- kragen 1y agoI just wrote a quick HTML renderer in Python with ReportLab: https://GitHub.com/kragen/dercuano/blob/master/genpdf.py https://GitHub.com/kragen/dercuano/blob/master/genpdf.py It only handles like 5% of HTML, but it's the 5% I was using. I've also had success producing PDFs with GhostScript from a PostScript file. PostScript is really easy to write, almost like SVG.
- Koffiepoeder 1y agoIf you want lots of differently styled templates, template management and editing/styling capabilites in word or excel (ie. you can just ask your customer/employer/.. to make an example document), I can really recommend Carbone [0]. I've been a happy customer for a few years now. Extra advantage is also that it also offers you excel outout generation as well, which is also often a requirement in applications. They have a SaaS offering as well if you'd like. They are open source though, so you can easily run a docker container! [0]: https://carbone.io/ https://carbone.io/
- crsr 1y agoBrowserless have browsers as a service, and a dedicated pdf endpoint[0] you can call. Had really good experience with this. [0] http://docs.browserless.io/rest-apis/pdf-api http://docs.browserless.io/rest-apis/pdf-api
- Animats 1y agoPrint to PDF in the browser? My main use for that is printing appointment information, tickets, and product listings. The product listings are useful when trying to find in a store something that's supposedly available and in stock. Usually, only the first page is useful. There will be additional useless pages of irrelevant items, deals, and ads.
- fredguth 1y agoI would use pandoc and convert to pdf using typst: ``` pandoc input.html -t typst -o output.typ typst compile output.typ output.pdf ```
- trollbridge 1y agoI wrote a solution in 2010 that used headless Firefox with some plugins to generate a PDF and then had the graphic designer write print CSSes. It was driven by Perl and was a convenient way for non-programmers to design forms. Unfortunately, that server and software stack is still around and still in production.
- znpy 1y ago> Unfortunately, that server and software stack is still around and still in production. that means you did a good job.
- Dwedit 1y ago2010-era Firefox is probably plagued by security holes.
- detaro 1y agoIf your print-file generation code tries to exploit the headless browser you use to turn its outputs into PDF something has gone very wrong already.
- trollbridge 1y agoMy biggest concern would be the Perl libraries I used to sanitise the input. I checked and none of them have any CVEs, though.
- busymom0 1y agoIf you are able to do this on a Mac, you can load the html in a WKWebView and then use the function: createPDF(configuration:completionHandler:) https://developer.apple.com/documentation/webkit/wkwebview/createpdf(configuration:completionhandler https://developer.apple.com/documentation/webkit/wkwebview/c...:)
- estimator7292 1y agoIIRC LibreOffice has some command line tools to do all kinds of document conversion
- freeopinion 1y agoClearly,a million people have tried to find an answer for this question. I've tried. At least one of my attempts was an XY problem. I was converting generated HTML that would never see a browser. It was never intended to see a browser. The people generating it were very good at HTML/CSS/JS, but didn't know how to produce the same content outside HTML.
- gangtao 1y agoI use chrome and ctrl+P
- jvanveen 1y agoPuppeteer (Pdfium => https://github.com/chromium/pdfium https://github.com/chromium/pdfium)
- stared 1y agoWhat's your goal? Print it? Archivize it? Send it via email? Read it on another device (which)? Depending on that, there are different solutions and trade-offs. For example on how to deal with pagination.
- Towaway69 1y agoOP wants to archive 5k a month --> https://news.ycombinator.com/item?id=45440073 https://news.ycombinator.com/item?id=45440073
- 0xMohan 1y agoOpen the html file in firefox and `ctrl + p`
- dvcoolarun 1y agoI built this: https://github.com/dvcoolarun/web2pdf https://github.com/dvcoolarun/web2pdf — a CLI tool for converting web pages to PDFs, recently open-sourced after adding several new features. (Might be useful!) Not related to the thread, but if anyone is looking to hire a developer or knows of opportunities, I was recently let go and am actively searching. Any leads or feedback would be greatly appreciated. Sample PDF: https://drive.google.com/file/d/1n7M1TKOptSsYiibrbvV_Yojx53TK3k5E/view https://drive.google.com/file/d/1n7M1TKOptSsYiibrbvV_Yojx53T...
- mimi_007 1y agoYeah, wkhtmltopdf can be a pain, especially with modern CSS/JS-heavy pages. One option you could try is PDFBolt - you can design a template in HTML/CSS once, then simply pass the JSON data and template ID. It handles dynamic content and modern layouts without all the quirks of wkhtmltopdf. You can also convert HTML content or URLs to PDFs.
- dredmorbius 1y agoShortcutting much of the discussion here (what are you goals / why would you do that / don't use format X): a key problem is that neither HTML (as published on today's Web) nor PDF are reliable as canonical document formats. Tagged-markup such as Markdown (or otherlightweight markup languages) or LaTeX (or other heavy markup languages) are far more robust. Markdown has its variants, but all are pretty simple and easy to produce. LaTeX is slightly more complex, but remains quite straightforward for simple works. Once you've got an appropriate canonical version in any of these options, you have an embarassment of riches to convert to any given document format (what I call endpoints) you'd care for: PDF, HTML, RTF, DOCX, or many, many others. I generally reach for Pandoc first, which itself, yes, of course, often relies on additional tools/libraries to parse or generate endpoints, but is quite versatile. You can simplify the intake of HTML by stripping out cruft. Readability, Beautiful Soup, or other HTML filtering tools can target the core content and metadata you most likely want. Otherwise, think through what you're doing and why to more narrowly define your goals and tools. E.g., if you want a faithful printed representation of a mainstream-browser-rendered page (that is, Google Chrome), you'd probably do best to use its print-to-PDF options (mentioned several times here). If you want to extract core text, filtering out much of today's WWW cruft will be a high priority.
- carlosjobim 1y agoOrion browser and export to PDF function. It will export the page exactly as it looks, including the page dimension.
- deafpolygon 1y agoPandoc or use a browser to save as pdf
- bevstratov 1y agoGotenberg https://gotenberg.dev/docs/routes https://gotenberg.dev/docs/routes We use it in production for pdf exports and reports generation. Just spin up a docker container and use a client library or REST API to send html data.