3 ms·
I am still looking for a good open source templating system that does use PDF files as input, so you are not bound to a special tool / format to generate the te
by POPOSYS 4y ago
I am still looking for a good open source templating system that does use PDF files as input, so you are not bound to a special tool / format to generate the templates but can use whatever produces PDF files itself. Does anybody know one?
All printing workflows use proprietary formats as input and bind you to one tool or producer. Could Scribus help with that?
- mrazomor 4y agoOne thing that worked well in practice (years ago): - design svg template (eg. in inkscape) - set placeholders (as svg is just xml -- {%address%} etc) - replace the placeholders with the actual value & use headless inkscape to produce PDF
- POPOSYS 4y agoThis is interesting. Does exist some kind of extraction of the PDF generating part of inkscape, so you do not have to start inkscape all over again for generating PDF files? Or is there a good and fast command line tool, that can produce PDF from Inkscape PDF files?
- tkp 4y agoInkscape can be used from the commandline to export pdfs : https://wiki.inkscape.org/wiki/Using_the_Command_Line https://wiki.inkscape.org/wiki/Using_the_Command_Line
- franga2000 4y agoInkscape's format is just SVG, so any tool that can do SVG-PDF will do. I've seen cairosvg, imagemagick and rsvg used for this, but if your deployment allows, it does make the most sense to use inkscape from the CLI to avoid any inconsistencies: $ inkscape rendered.svg --export-pdf=output.pdf I think there's also a way to do batch processing where you don't need to spin up a new inkscape process for each file (which takes time), but I don't remember how it works anymore.
- Archelaos 4y agoI am using XML ===XSLT===> LaTeX ===pdflatex/lualatex===> PDF for more than a decade now. The whole pipeline is driven by a batch file that takes all XML files in an input folder, uses a temp folder for the intermediate LaTeX, and an output folder for PDF. I can produce HTML files from the same XML sources directly: XML ===XSLT===> HTML. For differences between PDF and HTML versions I have some special tags and attributes in my XML sources. If I want to change something in the layout, I modify the XSLT script and run the old XML sources through the pipeline again in one go. There was some up front effort in designing the XML tag system and writing the XSLT scripts. But since my later layout changes were minor, the required tweaks were easy.
- froh 4y agowhich XML did you go with? docbook? Dita? some homegrown?
- Archelaos 4y agoA quite simple homegrown DTD. Making it as simple as possible keeps the complexity of the XSLT scripts low. Most of it is similar to HTML (<paragraph>, <italics>, <bold>) with a few special attributes to add some semantics or processing hints.[1] Sometimes I am using several intermediated steps to produce the final HTML. It might then be useful to version the DTDs with a fixed REQUIRED version attribute in the root element, that must to appear in the XML files, to avoid applying a wrong XSLT script to an outdated version of my XML sources.[2] For a customer who needed to import large semi-structured legacy Word documents from another company into a database system, I once implemented the following process: The Word documents were converted to a relatively simple homegrown XML format based on the structural elements of the Word document. The resulting XML documents were manually corrected where the structural elements were incorrect. Some special attributes were added inside the XML documents to associate text passages with already existing database keys. When this was finished, an XSLT script was applied that split the large XML file into smaller ones based on this database keys; a human readable prefix, the key and a date went into the file name. These files were converted in bulk to LaTeX and then to PDF. Afterwards, I used a little tool to bulk upload only the fresh PDFs into the correct database entries based on the keys in their filenames. For one of my side-projects, a C# application, I am using another, object-oriented approach, where I have an abstract base class for reporting and two derived classes, one that outputs HTML and one that outputs LaTeX. The LaTeX output is then fed into lualatex to produce PDFs. You can check out the free Herodotus edition of my (closed-source) Factonaut project at https://www.factonaut.com/ https://www.factonaut.com/ to see it in action. [1] Using parameter entities for re-usability, such as <!ENTITY % output_attr SYSTEM "output_attr.ent"> <!ATTLIST foo %output_attr; > <!ATTLIST bar %output_attr; > in the DTD referring to an `output_attr.ent` file with the following contents: output (pdfonly|htmlonly|all) "all" [2] The declaration in the DTD looks like: <!ATTLIST root version (1.0) #REQUIRED > and the XML must then look like: <root version="1.0"> ... </root>
- tkp 4y agoDepending on your needs, Scribus might help with it's scripting API[1], see ScribusGenerator[2] for example. Scribus and inkscape can import pdfs but it would be preferable to use a clean "source" format, pdf is AFAIK meant for output, like lossy compressed images. [1] https://impagina.org/scribus-scripter-api/ https://impagina.org/scribus-scripter-api/ [2] https://berteh.github.io/ScribusGenerator/ https://berteh.github.io/ScribusGenerator/
- jcynix 4y agoThis Perl module can read and modify PDFs: https://metacpan.org/pod/PDF::API2 https://metacpan.org/pod/PDF::API2
- viraptor 4y ago> that does use PDF files as input That's not a great idea. PDFs are a very "final presentation" format. They don't even really have a strict concept of a block of text. There are even tools which will happily put the separate letters is various places and call it a day. I've done something like that (just inserting short content + signature into a PDF) and do not want to touch PDF editing with a 10ft pole ever again. To preserve sanity, use a different format that actually understands the structure of the content, one step before the PDF is generated.
- POPOSYS 4y agoThank you very much for the hints!
- gettalong 4y agoThe problem with having a PDF file as template is that it gets very hard to define how new pages should be created whenever there is too much content to fit on the existing pages. E.g. line items in an invoice template. What you can do is to use a PDF file for all the static content that appears on each page.
- mpweiher 4y agoI used to make a tool that did this with Postscript and embedded EPSes, so you could use Quark or InDesign to create the templates. The problem with using PDF is that most tools flatten the output, so that makes things a bit difficult.