9 ms·
Please let this be capable of generating PDFs from HTML from the command line.
by masterleep 9y ago
Please let this be capable of generating PDFs from HTML from the command line.
- zwerdlds 9y agoFWIW, you can do that now, using phantomjs (which is chrome) http://phantomjs.org/screen-capture.html http://phantomjs.org/screen-capture.html
- velodrome 9y ago> FWIW, you can do that now, using phantomjs (which is chrome) http://phantomjs.org/screen-capture.html http://phantomjs.org/screen-capture.html It's based on webkit.
- theandrewbailey 9y agoOnce upon a time, but not anymore. https://www.chromium.org/blink https://www.chromium.org/blink
- deleted 9y ago[deleted]
- rubber_duck 9y agoPhantomJS is based on webkit (and an outdated version IIRC + a shitty JS interpreter)
- chews 9y agospecifically it's qtwebkit and an old version of it.
- eppsilon 9y agoThey recently released a beta version based on a newer version of QtWebKit: https://bitbucket.org/ariya/phantomjs/downloads/ https://bitbucket.org/ariya/phantomjs/downloads/
- dasmoth 9y agoIt's a useful tool (and huge thanks to those who built it -- and SlimerJS for that matter) but whenever I've reached for there's always been some issue to resolve, generally relating to the exact version and/or set of APIs supported. Headless mode (with PDF support -- which it looks like the latest version of the Chromium remote protocol does indeed have) built into a mainstream browsers is nearly guaranteed to be a smoother experience.
- chickenfries 9y agoAlso with https://wkhtmltopdf.org/ https://wkhtmltopdf.org/ and http://pandoc.org/ http://pandoc.org/
- pharrlax 9y agoTried all of these and haven't found any as good as http://www.nightmarejs.org/ http://www.nightmarejs.org/
- erichurkman 9y agoAnd, if you need much higher fidelity and control of HTML/CSS -> PDF, there's the fantastic Prince library, http://princexml.com/ http://princexml.com/ (nonfree) (I've been using Prince for over a decade, rendering everything from prescription labels, packing slips, receipts, resumes, books, and more. It's great.)
- jamespaden 9y agoThat's also an API service, http://docraptor.com http://docraptor.com, with a different pricing model that uses the Prince library for PDF rendering.
- zwerdlds 9y agoTough crowd.
- vvoyer 9y agoYes: https://chromedevtools.github.io/debugger-protocol-viewer/tot/Page/#method-printToPDF https://chromedevtools.github.io/debugger-protocol-viewer/to...
- jzfeng 9y agoThe print to pdf command currently only support default print settings. Adding support for customized page size, header and footer, dpi, etc. is in progress. Please see bug: https://bugs.chromium.org/p/chromium/issues/detail?id=603559 https://bugs.chromium.org/p/chromium/issues/detail?id=603559 for updates.
- masterleep 9y agoExcellent news, thanks!
- michael_miller 9y agoAny idea if you can configure the page DPI with this API?
- kenshaw 9y agoYes, the emulation domain has APIs for changing that.
- erikig 9y agohttps://wkhtmltopdf.org/ https://wkhtmltopdf.org/ You can download the two tools wkhtmltopdf and wkhtmltoimage which use WebKit to generate pdfs/images.
- uptown 9y agoNot sure of your exact use-case, but mPDF does this well. https://github.com/mpdf/mpdf https://github.com/mpdf/mpdf
- pizza 9y agoIs there an easy way to do the opposite? e.g. read PDFs as HTML (intended to be read through a remote shell), or text?
- jevinskie 9y agoYou might be able to tweak this pdf.js example to dump out its canvas element after it renders. https://github.com/mozilla/pdf.js/blob/master/examples/learning/helloworld64.html https://github.com/mozilla/pdf.js/blob/master/examples/learn...
- kccqzy 9y agoExtracting text from PDF is not hard, though PDF only contains low-level formatting instructions so the result might not be nice, especially if the original PDF has any non-trivial formatting, like pull quotes, multi-column text, etc. If you don't care about that or the correct "flow" of text, it should be easy enough to just find all the Tj and TJ operators and extract their operands. You might also need to reverse some ligatures though. Producing nice semantic HTML is much harder, though also easy if you don't mind every word in a separate absolutely positioned div. Many PDF reader software already contains empirically tuned routines to infer the text flow and generate text files (because the software needs to handle Select All and Copy), but they often produce bad results. But if you just want to read a PDF on a remote machine over ssh, the easiest solution might be just transferring the file and then opening it locally, or use X forwarding and open the PDF with a graphic reader.
- desdiv 9y agoI'm in the same boat. You might want to check these out: http://coolwanglu.github.io/pdf2htmlEX/ http://coolwanglu.github.io/pdf2htmlEX/ https://github.com/JonathanLink/PDFLayoutTextStripper https://github.com/JonathanLink/PDFLayoutTextStripper http://tabula.technology/ http://tabula.technology/ https://docparser.com https://docparser.com
- ufmace 9y agoEasy, not so much, depending on exactly what you want to get out of it. I did a project with this once https://www.idrsolutions.com/jpdf2html5/ https://www.idrsolutions.com/jpdf2html5/. Last I checked, they only supported it as a Java library that could generate rather nice looking and complete HTML pages from PDF documents. The output was great, but it was kinda pricey and difficult to work with. On the opposite side of the complexity level, I have also used this http://www.pdfsharp.com/PDFsharp/ http://www.pdfsharp.com/PDFsharp/ to extract bits of text from PDFs. It's free, but you only get access to the raw PDF text with formatting codes. It works fine if you just want to grab a short string, but you got your work cut out for you if you want to do anything more sophisticated.
- mixu 9y agoIf you're willing to wait until Electron releases a Chrome 59 -based build, I'll be updating https://github.com/mixu/electroshot https://github.com/mixu/electroshot which handles screenshots and print-to-PDF along with a bunch of other niceties.
- deleted 9y ago[deleted]
- frik 9y agoEspecially it should adhere to CSS3 page break properties.