3 ms·
It seems like it should be doable to train a two-tower model, or something similar, that simultaneously runs OCR on the image and tries to read through the raw
by oddthink 6y ago
It seems like it should be doable to train a two-tower model, or something similar, that simultaneously runs OCR on the image and tries to read through the raw PDF, that should be able to use the PDF to improve the results of the OCR.
Does anyone know of any attempt at this?
Blah blah blah transformer something BERT handwave handwave. I should ask the research folks. :-)
- pas 6y agoYes. But. The main problem is that in many cases the visual structure is some highly custom form that is hard to present to the user in text. And on top of this many times there are no text data in the PDF just a JBIG2 image per page.
- mcswell 6y agoI've thought about it, but haven't tried. As an experiment, we once tried converting an OCRed dictionary (this one: https://www.sil.org/resources/archives/10969 https://www.sil.org/resources/archives/10969) into an XML dictionary database. (There are probably better ways to get an XML version of that particular dictionary, but as I say, this was an experiment.) Despite the fact that it's a clean PDF, and uses a Latin script whose characters are quite similar to Spanish (and the glosses are in Spanish), the OCR was a major cause of problems: Treating the upside down exclamation as an 'i', failing to separate kerned characters, confusion between '1' and 'l', misinterpreting accented characters, and so on and so on. And for some reason the OCR was completely unable to distinguish bold from normal text, even though a human could do so standing several feet away. So I did think of extracting the characters from the PDF. If it had been a real use case, instead of an experiment, I might have done so. Write-up here: https://www.aclweb.org/anthology/W17-0112/ https://www.aclweb.org/anthology/W17-0112/
- oddthink 6y agoInteresting! We're working on OCR on menu photos, which has some parallels in structure, but has a much smaller common vocabulary than a dictionary, almost by necessity. :-) Many menus are also available in PDF form, so we're trying to figure out if it's worth bothering with the PDF itself, or if we should just render to image and thus reduce the problem to the menu-photo one.