6 ms·
ISO PDF spec is getting Brotli – ~20 % smaller documents with no quality loss
- delfinom 8mo agotl;dr Commerical entity is paying to have the ISO altered to "legalize" their SDK they are pushing which is incompatible with standard PDF readers. ISO is pay to play so :shrug:
- bhouston 8mo agoI'm no fan of Adobe, but it is not that hard to add brotli support given that it is open. Probably can be added by AI without much difficulty - it is a simple feature. I think compared to the ton of other complex features PDF has, this is an easy one.
- lmz 8mo agoIt's not even clear that they were the ones suggesting inclusion. They're just saying their library now supports the new thing. https://pdfa.org/brotli-compression-coming-to-pdf/ https://pdfa.org/brotli-compression-coming-to-pdf/ > As of March 2025, the current development version of MuPDF now supports reading PDF files with Brotli compression. The source is available from github.com/ArtifexSoftware/mupdf, and will be included as an experimental feature in the upcoming 1.26.0 release. > Similarly, the latest development version of Ghostscript can now read PDF files with Brotli compression. File creation functionality is underway. The next official Ghostscript release is scheduled for August this year, but the source is available now from github.com/ArtifexSoftware/Ghostpdl.
- adrian_b 8mo agoYes, I do not see any source of financial gain that could motivate them for this, because both MuPDF and Ghostscript are free. MuPDF is an excellent PDF reader, the fastest that I have ever tested. There are plenty of big PDF files where most other readers are annoyingly slow. It is my default PDF and EPUB reader, except that in very rare cases I encounter PDF files which MuPDF cannot understand, when I use other PDF readers (e.g. Okular).
- whizzx 8mo agoNo this feature is coming straight from the PDF association itself and we just added experimental support before it's officially in the spec to help testing between different sdk processors. So your comment is a falsehood
- bhouston 8mo agoAre they using a custom dictionary with Brotli designed for PDFs? I am not sure if it would help or not, but it seems like one of those cases it may help? Something like this: https://developer.chrome.com/blog/shared-dictionary-compression https://developer.chrome.com/blog/shared-dictionary-compress... In my applications, in the area of 3D, I've been moving away from Brotli because it is just so slow for large files. I prefer zstd, because it is like 10x faster for both compression and decompression.
- whizzx 8mo agoThe pdf association is still running experiments on whether or not to support custom dictionaries based on real life workloads gains. So it might land in the spec once it has proven if offers enough value
- Proclus 8mo agoIt seems they're using the standard dictionary, which is utterly bizzare. The standard Brotli dictionary bakes in a ton of assumptions about what the Web looked like in 2015, including not just which HTML tags were particularly common but also such things as which swear words were trendy. It doesn't seem reasonable to think that PDFs have symbol probabilities remotely similar to the web corpus Google used to come up with that dictionary. On top of that, it seems utterly daft to be baking that into a format which is expected to fit archival use cases and thus impose that 2015 dictionary on PDF readers for a century to come. I too would strongly prefer that they use zstd.
- bhouston 8mo agoBTW I've looked into custom dictionaries before for similar use cases and I suspect it would only offer like a 1% improvement or so for PDFs -- still good, but not a massive difference maker. The issue is that PDFs, like web pages, are incredibly repetitive in terms of their tags/structure. As such the custom dictionary only helps if the doc is really small, otherwise because of the repetitive nature, the self-inferred dictionary will resemble the custom dictionary after just a few blocks of PDF content. The sole exception is if they are restarting the brotli stream for each page, and they are not sharing a dictionary, custom or inferred across the whole doc. Then the dictionary will have to be re-inferred on each page, and then a shared custom dictionary would make more sense.
- bobpaw 8mo agoHow can iText claim that adding Brotli is not a backward incompatible change (in the "Why keep encoding separate" table)? In the first section the author states that any new feature must work seamlessly with existing readers. New documents created that include this compression would be unintelligible to any reader that only supports Deflate. Am I missing something? Adoption will take a long time if you can't be confident the receiver of a document or viewers of a publication will be able to open the file.
- whizzx 8mo agoIt's prototypish work to support it before it land's in the official specification. But it will indeed take some adoption time. Because I'm doing the work to patch in support across different viewers to help adoption grow. And once the big opensource ones ship it pdfjs, poppler, pdfium, adoption can quickly rise.
- croes 8mo agoThere are old devices where the viewer can’t be patched. That’s killing one of the main features of PDF
- ericpauley 8mo agoSome real cognitive dissonance in this article… “The PDF Association operates under a strict principle—any new feature must work seamlessly with existing readers” followed by introducing compression as a breaking change in the same paragraph. All this for brotli… on a read-many format like pdf zstd’s decompression speed is a much better fit.
- xxs 8mo agoyup, zstd is better. Overall use zstd for pretty much anything that can benefit from a general purpose compression. It's a beyond excellent library, tool, and an algorithm (set of). Brotli w/o a custom dictionary is a weird choice to begin with.
- greenavocado 8mo agoThis bizzare move has all the hallmarks of embrace-extend-extinguish rather than technical excellence
- adzm 8mo agoBrotli makes a bit of sense considering this is a static asset; it compresses somewhat more than zstd. This is why brotli is pretty ubiquitous for precompressed static assets on the Web. That said, I personally prefer zstd as well, it's been a great general use lib.
- dist-epoch 8mo agoYou need to crank up zstd compression level. zstd is Pareto better than brotli - compresses better and faster
- jeffbee 8mo agoAre you sure? Admittedly I only have 1 PDF in my homedir, but no combination of flags to zstd gets it to match the size of brotli's output on that particular file. Even zstd --long --ultra -22.
- ksec 8mo agoWhy not zstd?
- PunchyHamster 8mo agoincompetence
- deleted 8mo ago[deleted]
- whizzx 8mo agoYou can read about it here https://pdfa.org/brotli-compression-coming-to-pdf/ https://pdfa.org/brotli-compression-coming-to-pdf/
- jeffbee 8mo agoThat mentions zstd in a weird incomplete sentence, but never compares it.
- F3nd0 8mo agoThey don’t seem to provide a detailed comparison showing how each compression scheme fared at every task, but they do list (some of) their criteria and say they found Brotli the best of the bunch. I can’t tell if that’s a sensible conclusion or not, though. Maybe Brotli did better on code size or memory use?
- eviks 8mo agoHey, they did all the work and more, trust them!!! > Experts in the PDF Association’s PDF TWG undertook theoretical and experimental analysis of these schemes, reviewing decompression speed, compression speed, compression ratio achieved, memory usage, code size, standardisation, IP, interoperability, prototyping, sample file creation, and other due diligence tasks.
- 8mo ago
- cess11 8mo ago'Your PDF:s will open slower because we decided that the CDN providers are more important than you'. If size was important to users then it wouldn't be so common that systems providers crap out huge PDF files consisting mainly of layout junk 'sophistication' with rounded borders and whatnot. The PDF/A stuff I've built stays under 1 MB for hundreds of pages of information, because it's text placed in a typographically sensible manner.
- h4x0rr 8mo agoWouldn't lzma2 be better here since a pdf is more read heavy?
- F3nd0 8mo agoGoing by one of Brotli’s authors’ comment [1] on another post, it probably wouldn’t. [1] https://news.ycombinator.com/item?id=46035817 https://news.ycombinator.com/item?id=46035817
- nialse 8mo agoWho is responsible for the terrible decision? In the pro vs con analysis, saving 20% size occasionally vs updating ALL pdf libraries/apps/viewers ever built SHOULD be a no-brainer.
- ndriscoll 8mo agoWhat is the point of using a generic compression algorithm in a file format? Does this actually get you much over turning on filesystem and transport compression, which can transparently swap the generic algorithm (e.g. my files are already all zstd compressed. HTTP can already negotiate brotli or zstd)? If it's not tuned to the application, it seems like it's better to leave it uncompressed and let the user decide what they want (e.g. people noting tradeoffs with bro vs zstd; let the person who has to live with the tradeoff decide it, not the original file author).
- eru 8mo agoWell, if sanity had prevailed, we would have likely stuck to .ps.gz (or you favourite compression format), instead of ending up with PDF. Though we might still want to restrict the subset of PostScript that we allow. The full language might be a bit too general to take from untrusted third parties.
- dunham 8mo agoDon't you end up with PDF if you start with PS and restrict it to a subset? And maybe normalize the structure of the file a little. The structure is nice when you want to take the content and draw a bit more on the page. Or when subsetting/combining files. I suspect PDF was fairly sane in the initial incarnation, and it's the extra garbage that they've added since then that is a source of pain. I'm not a big fan of this additional change (nor any of the javascript/etc), but I would be fine with people leaving content streams uncompressed and running the whole file through brotli or something.
- mikkupikku 8mo agoI thought PDFs can contain arbitrary PS.
- eru 8mo ago> Don't you end up with PDF if you start with PS and restrict it to a subset? PDF is also a binary format.
- avalys 8mo agoThis article is AI slop.
- jeffbee 8mo agoYep.
- vgtftf 8mo ago[flagged]
- superkuh 8mo agoThis is nice, but PDF jumped the shark already. It's no longer a document format that always looks the same everywhere. The inclusion of "Dynamic XFA (XML Form Architecture) PDF" in the spec made it so PDF is an unreliable format. The aformentioned is a PDF without content that pulls down all it's content from the web. It even still, ostensibly, supports Flash (swf) animations. In practice these "PDF"s are just empty white pages with an error message like, >"Please wait... If this message is not eventually replaced by the proper contents of the document, your PDF viewer may not be able to display this type of document. You can upgrade to the latest version of Adobe Reader for Windows®, Mac, or Linux® by visiting http://www.adobe.com/go/reader_download http://www.adobe.com/go/reader_download. For more assistance with Adobe Reader visit http://www.adobe.com/go/acrreader http://www.adobe.com/go/acrreader. Windows is either a registered trademark or a trademark of Microsoft Corporation in the United States and/or other countries. Mac is a trademark of Apple Inc., registered in the United States and other countries. Linux is the registered trademark of Linus Torvalds in the U.S. and other countries."
- kayodelycaon 8mo agoFortunately, XFA is deprecated. I haven’t seen one of those for a very long time.
- superkuh 8mo agoMaybe in spec, but the damage is done and persists. The (USA) Wisconsin Dept. of Natural Resources has nearly all their regulation PDFs as these XFA non-pdfs that I cannot read. So I cannot know the regulations. My emails about this topic (to multiple addresses over many years a dozen times) have gone unanswered. If Acrobat supports it it doesn't matter what the spec says. Until Adobe drops XFA from Acrobat and forces these extremely silly people to stop, PDF is no longer PDF.
- whinvik 8mo agoI am often frustrated by PDF issues such as how complicated it is to create one. But reading the article I realized PDFs have become ubiquitous because of its insistence on backwards compatibility. Maybe for some things it's good to move this slow.
- jhealy 8mo agoThe article is wrong, the PDF spec has introduced breaking changes plenty of times. It’s done slowly and conservatively though, particularly now that the format is an ISO spec. The PDF format is versioned, and in the past new versions have introduced things like new types of encryption. It’s quite probable that a v1.7 compliant PDF won’t open on a reader app written when v1.3 was the latest standard.
- nbevans 8mo agoThis is a really really bad idea. Don't break backwards compat. for 20% of gains. Internet connection speeds and storage capacities only go up. In a few years time, 20% of gains will seem crazy to have broken back-compat for.
- gcr 8mo agoIf we're making breaking changes to PDFs, I'd love if the committee added a modern image format like JPEG-XL. In my experience, most disk usage of PDFs comes from images, not streams. I keep a bunch of comics in PDF but JPEG-XL is by far the best way to enjoy them in terms of disk space.
- Bolwin 8mo agoOdd you should say that, as that's exactly what they've been discussing
- gcr 8mo agoNo it's not. This article is about proposing Brotli as another possible '/Filter' for stream objects, like content streams (page drawing commands). Images are streams too, but unless you mean compressing raw pixel bytes in Brotli, there's no mention of a JPEG-XL or WEBP filter.
- NoahZuniga 8mo agowell, not mentioned in this specific article. But JPEG-XL support is something they're working on [1]. [1]: https://pdfa.org/wp-content/uploads/2025/10/PDFDays2025-BreakingBad-Wyatt.pdf https://pdfa.org/wp-content/uploads/2025/10/PDFDays2025-Brea...
- gcr 8mo agoOh cool!! TIL
- deleted 8mo ago[deleted]