9 ms·
The surprisingly complex journey to text-selectable client-side generated PDFs
- josefrichter 5mo agoIt’s not that surprising. It’s one of those well known pandora boxes of web development: email templates, PDFs, printing,…
- FailMore 5mo agoAh, I didn't know that. It's not something I had worked on before, and the file format is highly prevalent (so I assumed things would be easy), so it was surprising to me
- SirHumphrey 5mo agoNothing about PDF is easy. Similarly to what once Tom Scott said about time zones, every time I must deal with PDFs I pray that PDF.js can be hacked in to doing it instead, otherwise I just don’t bother. It’s on of the few examples when converting it in to picture and chucking it in a multimodal llm is a more sensible solution than trying to parse it.
- caspper69 5mo agoYou would think that, but PDF is not really a format for text. It's a format that describes typography and graphics layout & formatting. It's not uncommon for a text pdf to not contain all of the text it renders (due to ligatures).
- deleted 5mo ago[deleted]
- ashishb 5mo agoSoftware engineers drastically underestimates GUI - Web layouts, mobile app layouts, and even PDF layouts are non-trivial pieces of work to get right in all circumstances.
- FailMore 5mo agoYep, they (can) rarely enter your domain... so it's easy to assume its going to be trivial (maybe because things like .md or .txt files are trivial, so it's easy to think there's not much of a delta)
- freedomben 5mo agoNobody who has actually worked on those things think that. You might want to qualify if you're only talking about people who have never worked in this area. In my experience it's the NON software engineers who tend to underestimate the complexity
- ashishb 5mo agoI didn't qualify. And the reason was that majority of backend engineers have never worked on frontend. At almost all big companies, the team working on the frontend problem is relatively small +and that's how it should be). > In my experience it's the NON software engineers who tend to underestimate the complexity Yeah. That too as well.
- gobdovan 5mo agoThanks, this puts into perspective why copy-paste from PDFs is so bad. I months into building a pasteboard transform library that normalises VS Code, Google Docs, PDFs and a bunch of Chromium apps provider-specific data so I can start pasting everything everywhere exactly how I want it. It's much, much messier than I expected. Apps put different UTTypes on the pasteboard that are not really compatible with each other. Usually there's a plain text fallback, then rich text/HTML, then provider-specific data. You show how much insane work is needed just to make text selectable with glyph mappings, layout, links, code blocks, rendered styles, etc. But once you copy from that PDF, most viewers still only expose raw text, and often broken raw text at that...
- FailMore 5mo agoYep, it is a very interesting space for improvements imo. Kind of broadly speaking copy and paste is so central to working with a computer in a smooth way it should probably have more power / quality built into it (e.g. not having to install some random plug in to get clipboard history, etc.)
- gobdovan 5mo agoPart of it is security/privacy and providers avoiding liability. People constantly copy passwords, tokens, personal data, etc., so clipboard history is risky by default. Apple probably does not want to expose a rich API here and then be responsible for securing that surface forever. So macOS does not really give you a clean "this app copied this semantic object" API. Clipboard-history apps generally poll NSPasteboard.changeCount, which already makes provenance fuzzy, since you can observe that the pasteboard changed, but not reliably know the source app. Pasting is fuzzy too. You know what representations were available, but not what the destination app actually accepted, because that decision happens inside the app and is generally opaque for the OS. So what even is history? Is it the raw object, the fallback text, the richest representation, the thing you intended to paste, or the thing the target app consumed? And even if you define history as "the observed events", polling can also miss states. And once you add transforms (like I want to), you basically have to define your own history model. A coherent OS clipboard-history API probably will never happen without big effort and liability policy changes from providers.
- Worf 5mo agoPDFs should be only for printing or maybe for keeping scanned versions of things. For anything else they're just not the right tool for the job. Not for things meant to be accessed on a computer like books, scientific papers or, for some weird reason, catalogs and price lists from websites. We have responsive and open standards like HTML and EPUB (zipped XTML) and they work great. arXiv has HTML papers, and libgen and anna's archive often have EPUB versions of books. The issue for me with EPUB is the lack of good readers now.
- FailMore 5mo agoInteresting point. What do you feel about the "business world"'s heavy use of PDFs? There is something to be said about the file format being trusted/so dominant now... probably some random sequence of events led to this happening... but perhaps hard to shift
- Sharlin 5mo agoBecause the business world used to run on paper, and pdf became the de facto standard desktop publishing file format because Adobe became the de facto king of desktop publishing. Storing, transferring, and reading documents on paper has given way for doing all of that digitally, but path dependency guarantees that there’s no way of getting rid of PDF now. Purely psychologically, I think there’s something that feels more "secure" or long-lasting about PDF’s perceived quasi-immutability compared to formats designed to be edited.
- FailMore 5mo agoYeah, I think the point about editing is a very good one. There is something comforting about them and perhaps that's it (+ maybe we are used to them being A4 pages, so you know what to expect). I think also the lack of flexibility with rendering is good, if you see it on one device you know exactly how it will look on another device.
- Worf 5mo ago
- alansaber 5mo agoYou don't know the hell of trawling through PDF XML and HTML construction until you've done it
- cbolton 5mo agoI wonder if using Typst would be a viable solution: the compiler can be built into a wasm component that runs locally in the browser (that's what the Typst webapp does) and it generates good PDFs with working selection/copy/paste. There's even a package (cmarker) than can translate Markdown to Typst which could be enough for a MVP.
- LuttelBurchtje 5mo agoThe Swiss army knife for document conversion, Pandoc, supports compiling to WASM since 3.9 [1]. It supports Markdown, in a wide variety of flavours, and Typst. Their official demo page provides a PDF output via Typst, all done client-side [2]. Furthermore, you get .docx and other output formats as well [1]: https://github.com/jgm/pandoc/releases/tag/3.9 https://github.com/jgm/pandoc/releases/tag/3.9 [2]: https://pandoc.org/app/ https://pandoc.org/app/
- 000000000001 5mo agoI vibecoded a pdf replacement at work, sort of. I wanted a way to make submitting Inventory Changes at work easier, so I took the pdf, used StirlingPDF to convert to an html bundled zip, I converted the .png with the form border and symbols to base64, and then wrote a powershell script to replace <p> tags with variable data from a csv export of our inv data(I tried to use odbc to extract it but once a dev showed me the logical, physical, and views that made up our Inventory lookup I went back to using a xlsx export that is built into our environment and letting the ps1 trim and sanitatize the input). As the conversion places text with absolute positioning, I was able to fine tune the layout and spacing. I then used my local AI qwen3.6-27b to convert my ps1 to a single html binary webapp with html/css/js, no external framework, two js scripts are loaded via cdn for now. Inspired by how well that worked I vibecoded a drag and drop editor to build forms for other processes, I upload a png w which gets converted to base64 and then I can drag and drop text elements to where they need to be and export. I know how many people feel about AI coded projects so these are really only for me, I didn't expect my coworkers to adopt it or anything but they did.
- FailMore 5mo agoSounds pretty interesting. Got a video of it? Sometimes our vibe coded tools are pretty useful... but sometimes we can be a bit shy to share them given the vibeness...
- gadders 5mo agoWhile I think you would be stupid to try and vibe-code EG a banking platform or SAP, I do think there is a lot of scope for small apps that can solve a particular issue. I've just vibe-coded a mobile app, and I'm comfortable doing that because this particular app can't update data or send data anywhere. It's damage blast radius is pretty low.
- mvdwoord 5mo ago"everyone hates PDFs where you can't reliably select and copy text!" Boy do I. One of my biggest annoyances is receiving an invoice in pdf format, where I can either not select the text at all, or where you cannot cleanly select text, i.e. when you try to select something it somehow half highlights the line above as well and I am not sure what is on my clipboard, and need to paste temporarily in a text editor, then select what I need ... etc Super nice when the list IBAN numbers for payment in a tiny font size as well. Maybe I should vibecode a little helper. tool to visually select a rectangle and perform OCR and detect IBAN numbers or show a popup with proper text to do my subselect.
- 13324 5mo agoPersonal lifehack: Use the address bar of your browser to view the clipboard content quickly and to omit any formatting from it.
- cluckindan 5mo agoThis also conveniently sends it to your search provider, and possibly to the browser vendor for analytics.
- Steve44 5mo agoWe have a couple of large customers who will only send remittance advices as a PDF, the are several pages and a couple of hundred rows. Apparently their system can not send XLSX or any other format. I've been a happy user of Tabula[1] for a few years and it works really well, for my needs anyway. I just import, auto-detect tables, select "Stream", and then export to a CSV. [1] https://tabula.technology/ https://tabula.technology/ [1] https://github.com/tabulapdf/tabula https://github.com/tabulapdf/tabula
- yolo3000 5mo agoToday I had to edit in MacOS a pdf which had some text fields. It had 3 places which resembled checkboxes, so in those I was able to place an `x` character. When I saved it and previewed it, the `x` was in 2 places, but not the third. I tried several times but couldn't make it work. I gave up, started a Claude session, and asked it to fill in the `x`. 4k tokens of Python later it managed to place the `x` correctly.
- lavrton 5mo agoJust was solving the exact same issue. Recently I released https://polotno.com/render-tag/ https://polotno.com/render-tag/ library to render rich text into 2d canvas context. And it turns out it was very easy to adapt it to work with pdflib library (via 2d canvas <-> pdf context) proxy. I was able to render good set of rich text features. Thinking to make that bridge open source as well. Maybe you will be interested in that?
- FailMore 5mo agoYep sounds good!
- thomas_viaelo 5mo ago[dead]
- memeks 5mo ago[dead]
- ak217 5mo agoTangentially related, one of the most underappreciated projects for print media and PDF generation is paged.js. It goes down the rabbit hole of paged media and the surprising complexity of it (have you ever thought about what it takes to render a table across multiple pages?) and provides a great foundation for solving these problems with sanity and using open web standards. It's a project that deserves more support.
- FailMore 5mo agoNice shout - I will check it out - this I assume: https://github.com/pagedjs/pagedjs/ https://github.com/pagedjs/pagedjs/
- jiehong 5mo agoThis page made me discover the tool, and I find sdocs.dev pretty cool!
- FailMore 5mo agoThank you very much! I launched via a Show HN a few weeks ago [1] which goes into quite a few details which might not be obvious. Also one thing that probably isn't clear with everything I've written so far is that by far the most enjoyable use case I've found with SmallDocs (or the `sdoc` cli command) is telling your CLI-based agent to "sdoc" you things. E.g.: * "write up the plan and sdoc it to me" * "explain async/await to me simply in a sdoc" * "draft the release notes as a sdoc I can send to Ben for feedback" * etc. [1] https://news.ycombinator.com/item?id=47777633 https://news.ycombinator.com/item?id=47777633
- evolve2k 5mo agoThe second part of the article moves into a very confusing to read user experience. That said, maybe I missed it, but I didn’t see any mention of pandoc that is known to do markdown > pdf rendering “client side”. This wasn’t AI written by chance?
- pimlottc 5mo agoI would look for any other well-known formats that already have a dependable pipeline for rendering to a selectable text PDF. Perhaps LaTeX? Then make a converter to covert from markdown to that format.
- zameermfm 5mo agoThere was a plug of smalldocs intro in that post, regarding that, Dont we need the md files on the wikis and documentations in the central project management tools, instead of client side anything?
- FailMore 5mo agoSmallDocs person here, IMO the location of Markdown files and the reading of them are separate things. SmallDocs is just about reading not about storing. Eg if you are working in ~/code/my_project, and have a README.md, you (or your agent) can `sdoc README.md`. This opens up the file for reading (in the browser, but 100% privately). It doesn’t change the stored location of the Markdown file. Is that what you were getting at or something else?
- zameermfm 5mo agoAh its a reader, okay, great tool!