4 ms·
Let me assure you, this does exist in the wild. I spotted it many times, especially in documents scanned 15+ years ago, when space and bandwidth were an issue a
by monista 8y ago
Let me assure you, this does exist in the wild. I spotted it many times, especially in documents scanned 15+ years ago, when space and bandwidth were an issue and people did anything to reduce file size, at the cost of quality. It was rare enough to not make the format unusable, but not just once, making you worry "was it 3 sp. or 8 sp.?"
It seems also that the problem happens more often with cyrillic glyphs which have less descending/accending elements that typical latin ones.
- jacquesm 8y agoWith DjVu or with ocr'd documents in general?
- monista 8y agoWhile of course it's an issue with any scanned docs, the DjVu compression method makes it harsher, as, say, you can end with _all_ '8's in document replaced with '3' (or, say, 'h's with 'n's), at the same time keeping seemingly decent appearance so that you don't suspect that something is wrong.
- jacquesm 8y agoI'd love to see some concrete examples. I totally understand how this could theoretically happen but it all hinges on it actually happening in the documents that I'm working with and to have a test case where it happens for sure would make all the difference in determining whether or not the documents I'm working with are at risk or not and if so if I can detect which ones are at risk (that would be half the battle won). It would have to be a case where a human would see a 3 or an 8 (or a similar transposition with other glyphs) resulting in a document that is corrupted in such a way that afterwards the human would see the (invalid) alternative. Even a single instance of this happening would be very relevant.
- monista 8y agoWell I know that there are examples in some of 8k+ djvu files on my hard disk (mostly old textbooks), as I learned about this problem many years ago from live experience. But I couldn't invent a method to find an example deliberately, so you can only take my word that they exist. To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. The compression algorithm could not only mistake '3' as '8' with small gaps left, it would replace this dirty '3' with the image of reference '8' image, so that human reader think that he sees a scan, i.e. an image of page scanned, while in fact it sees 'edited' image, not corresponding to actual page.
- jacquesm 8y ago> so you can only take my word that they exist. I take your word, there is a technical reason behind this, not that I don't believe you. > To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. Yes, I totally get that. It's about DjVu's compressor replacing the image of one character with the image of another.
- andrewshadura 8y agoIIRC Xerox revoked lots of their copiers which used the same compression algorithm internally after this was demonstrated by a document with lots of numbers. Google it, it's not that hard to find (I'm on a phone now so it's not very convenient for me atm)
- LeoPanthera 8y ago> Let me assure you, this does exist in the wild. You would obtain more credibility by actually providing such an example.