3 ms·
> I believe it was spitting out perfectly formed zeroes where another digit was in the original. Yikes. JBIG2 lossy compression. Covered in another hn story: h
by ElectricalUnion 3y ago
> I believe it was spitting out perfectly formed zeroes where another digit was in the original. Yikes.
JBIG2 lossy compression. Covered in another hn story: https://news.ycombinator.com/item?id=29223815 https://news.ycombinator.com/item?id=29223815
From the story and the comments:
> This is not an OCR problem (as we switched off OCR on purpose), it is a lot worse – patches of the pixel data are randomly replaced in a very subtle and dangerous way: The scanned images look correct at first glance, even though numbers may actually be incorrect.
It's lossy compression turned to 11 that looks convincingly non-lossy.
- hinkley 3y agoI know technically you are right, but I feel like if you sat down and described the high level features and goals of that image compression (to compress text to a ridiculous degree) and OCR , you’d be hard pressed to say which list is which unless they gave it away with jargon words. My brain compressed it to “OCR” and filed it there. It’s like Jim Gaffigan’s joke about working in a TexMex restaurant in the northern Midwest. It’s a tortilla, cheese, meat and beans… I tell you what I’ll just bring you something.