3 ms·
I would start by either improving the OCR, or piggyback on GCV/Amzn, who have better tools. For instance, what if I drag-and-drop their sample image to this dem
by zzleeper 6y ago
I would start by either improving the OCR, or piggyback on GCV/Amzn, who have better tools. For instance, what if I drag-and-drop their sample image to this demo page https://cloud.google.com/vision https://cloud.google.com/vision ?
All text is recognized, and as far as I can tell, there are no errors.
There are a lot of difficulties with older text, but I would start with low-hanging fruit such as trying to use better tools. On top of that, you have the problem of making sense of the layout, fixing common typos, etc.
- est31 6y agoI tried the GCV demo as well and got one mistake "prevent:" instead of "prevents" as it is in the text. But that's the only I could find, which puts it several categories ahead of the txt file. I'm not sure though whether the IA can get Google to OCR it for them for a budget they can afford. Likely they'd want OCR solutions that have a one-time cost, so volume based SAAS offerings won't work.