3 ms·
Optical character recognition... Has been pretty widespread tech since ~2016 or so.
by vren 3y ago
Optical character recognition... Has been pretty widespread tech since ~2016 or so.
- hinkley 3y agoOh my ex and some acquaintances were bitching about it in 1999. I recall a guy with good transcription skills starting a long thread on… must have been slashdot? About how it was faster for him to transcribe the numbers again than to verify the OCR was correct, but he couldn’t tell his boss that and what should he do? There’s also a famous case where the compression algorithm in a copy/fax machine had a bug in a predefined dictionary that resulted in changing digits to other digits when making copies or faxes. Because they weren’t just compressing, they were trying to use a specialized OCR to do aggressive compression. I believe it was spitting out perfectly formed zeroes where another digit was in the original. Yikes.
- ipython 3y agoyou're referring to the Xerox Workcenter: http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...
- hinkley 3y agoYES. God what a nightmare. Can you imagine all of the arguments and threats of lawsuits that bug caused for small and medium businesses? You changed the contract after we had a verbal agreement and tricked me into signing! I wonder if any executive assistants got fired over that. Surely.
- ElectricalUnion 3y ago> I believe it was spitting out perfectly formed zeroes where another digit was in the original. Yikes. JBIG2 lossy compression. Covered in another hn story: https://news.ycombinator.com/item?id=29223815 https://news.ycombinator.com/item?id=29223815 From the story and the comments: > This is not an OCR problem (as we switched off OCR on purpose), it is a lot worse – patches of the pixel data are randomly replaced in a very subtle and dangerous way: The scanned images look correct at first glance, even though numbers may actually be incorrect. It's lossy compression turned to 11 that looks convincingly non-lossy.
- hinkley 3y agoI know technically you are right, but I feel like if you sat down and described the high level features and goals of that image compression (to compress text to a ridiculous degree) and OCR , you’d be hard pressed to say which list is which unless they gave it away with jargon words. My brain compressed it to “OCR” and filed it there. It’s like Jim Gaffigan’s joke about working in a TexMex restaurant in the northern Midwest. It’s a tortilla, cheese, meat and beans… I tell you what I’ll just bring you something.