2 ms·
Well I know that there are examples in some of 8k+ djvu files on my hard disk (mostly old textbooks), as I learned about this problem many years ago from live e
by monista 8y ago
Well I know that there are examples in some of 8k+ djvu files on my hard disk (mostly old textbooks), as I learned about this problem many years ago from live experience. But I couldn't invent a method to find an example deliberately, so you can only take my word that they exist.
To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc.
It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. The compression algorithm could not only mistake '3' as '8' with small gaps left, it would replace this dirty '3' with the image of reference '8' image, so that human reader think that he sees a scan, i.e. an image of page scanned, while in fact it sees 'edited' image, not corresponding to actual page.
- jacquesm 8y ago> so you can only take my word that they exist. I take your word, there is a technical reason behind this, not that I don't believe you. > To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. Yes, I totally get that. It's about DjVu's compressor replacing the image of one character with the image of another.