4 ms·
While of course it's an issue with any scanned docs, the DjVu compression method makes it harsher, as, say, you can end with _all_ '8's in document replaced wit
by monista 8y ago
While of course it's an issue with any scanned docs, the DjVu compression method makes it harsher, as, say, you can end with _all_ '8's in document replaced with '3' (or, say, 'h's with 'n's), at the same time keeping seemingly decent appearance so that you don't suspect that something is wrong.
- jacquesm 8y agoI'd love to see some concrete examples. I totally understand how this could theoretically happen but it all hinges on it actually happening in the documents that I'm working with and to have a test case where it happens for sure would make all the difference in determining whether or not the documents I'm working with are at risk or not and if so if I can detect which ones are at risk (that would be half the battle won). It would have to be a case where a human would see a 3 or an 8 (or a similar transposition with other glyphs) resulting in a document that is corrupted in such a way that afterwards the human would see the (invalid) alternative. Even a single instance of this happening would be very relevant.
- monista 8y agoWell I know that there are examples in some of 8k+ djvu files on my hard disk (mostly old textbooks), as I learned about this problem many years ago from live experience. But I couldn't invent a method to find an example deliberately, so you can only take my word that they exist. To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. The compression algorithm could not only mistake '3' as '8' with small gaps left, it would replace this dirty '3' with the image of reference '8' image, so that human reader think that he sees a scan, i.e. an image of page scanned, while in fact it sees 'edited' image, not corresponding to actual page.
- jacquesm 8y ago> so you can only take my word that they exist. I take your word, there is a technical reason behind this, not that I don't believe you. > To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. Yes, I totally get that. It's about DjVu's compressor replacing the image of one character with the image of another.
- andrewshadura 8y agoIIRC Xerox revoked lots of their copiers which used the same compression algorithm internally after this was demonstrated by a document with lots of numbers. Google it, it's not that hard to find (I'm on a phone now so it's not very convenient for me atm)