3 ms·
Strangely enough (?) I have not had many problems with PDF, certainly no more than with other document formats. I often use tools like pdftk [0] (as found on ma
by Yetanfou 6y ago
Strangely enough (?) I have not had many problems with PDF, certainly no more than with other document formats. I often use tools like pdftk [0] (as found on many Linux distributions) to split sections out of PDFs, create new ones out of single-page PDFs created with ImageMagick, create odd-even page versions etc. I generally do not touch the more "advanced features" like embedded JS (which I have disabled in all readers which support it), I just use it as a document format which more or less guarantees the resulting document looks the way it was meant to, plus or minus a few fonts. For that purpose it works works well enough.
The "bunch of jpgs/pngs/pick your poison in a tar container" format you describe exists in a fashion: Comic Book Archive, a format meant for and mostly used for comics. It consists of a compressed archive which contains sequentially numbered "pages" which can be JPEG, PNG or other image file formats. For pure image documents it can be used as a replacement for PDF but since it does not support text it can not be used for scanned OCR'ed documents. DjVu [2] does support a text layer but that comes at the cost of complexity, it is far from the simple container you propose. Since an OCR'ed text layer needs to save not only the text itself but also the location on the image for each character I don't see any way to avoid complexity here.
[0] https://www.pdflabs.com/tools/pdftk-server/ https://www.pdflabs.com/tools/pdftk-server/
[1] https://en.wikipedia.org/wiki/Comic_book_archive https://en.wikipedia.org/wiki/Comic_book_archive
[2] https://en.wikipedia.org/wiki/DjVu https://en.wikipedia.org/wiki/DjVu