4 ms·
my attempts at using AI to do OCR have always resulted in invented artifacts, which is not production feasible. does this suffer from that as well? A simple ex
by pmarreck 4mo ago
my attempts at using AI to do OCR have always resulted in invented artifacts, which is not production feasible. does this suffer from that as well?
A simple example is words that are supposed to be in other languages being automatically translated to English, which ruins the effect
- drakmo 4mo agoIf I would want to achieve 100% recognition results I would combine this method with an image model recreating the original document from the transcribed text and matching the layout. One can do that with using all but the page or paragraph from the document you want to recreate (to avoid recreating the exact passage under test from the image artifact directly). After reconstructing you can do an optical comparison that specifically matches misaligned characters and find the errors. Rinse and repeat. Expensive but it would guarantee 100% recognition.
- pbhjpbhj 4mo agoYou almost don't want [super-]word level ML (ie word-pair/phrase/sentence/document/corpus level). In transcription, you want near certainty, or you want marking that the word could not be read with certainty - yes, context lets you guess, but you want - for some OCR - to know when it's a guess based on other than the letters in order forming a word. Example, in a census document on familysearch.com the transcriber "corrected" a name as Joseph. The literal letters in the handwritten document spell Josepth ... and sure enough that's a local variant spelling (Eire). In another document the writer has used "Joh" as an abbreviation, a [human, I assume] transcriber put that as John ... which is most likely, but happens to be wrong. Sometimes you care that it's guessed, sometimes you want just the best guess.
- messe 4mo ago> Eire A nitpick, because it's often a dogwhistle: but almost nobody in Ireland calls it that when speaking English. And that's still incorrect in Irish, the correct spelling is Éire.
- pbhjpbhj 4mo agoBy saying it's a dogwhistle are you saying that not adding the correct diacritics is considered racist by Irish people? If I change the rest of the sentence to Na Gaeilge will that be better.
- messe 4mo agoNo, I'm saying that it's associated with a certain outdated and bigoted attitude toward the Irish. Using Éire in English, would be seen as odd. You wouldn't say Deutschland or Danmark. > If I change the rest of the sentence to Na Gaeilge will that be better. No. And you've used the genitive instead of nominative there, so I have some doubts that you could.
- pmarreck 4mo agoNot OC but couldn't help commenting here because I think this is a problem of subjectivity vs. intent and the ambiguity introduced by text. 1) I am German descent so I'd definitely use Deutschland to appear fancy or play with words when speaking of Germany, without any bias implied or meant. 2) The problem with believing in dogwhistles (whether they exist or not, and I know they do, but bear with me) is that the "perceived dogwhistle surface area" increases in proportion to your belief in the prevalence of dogwhistles. In other words, the more firmly you are looking for "plausibly deniable" racist terms, the more you will find terms that were actually intended to be innocent, to be offensive, and the more upset you will be in the world, AND the more annoyed people will get with you if they are not subscribed to the whole "we must avoid any possible term that could remotely be misconstrued as a plausibly-deniable dogwhistle for fear of offending someone" worldview. I would have absolutely used Éire but in a friendly way, and you're saying it would be perceived as a dogwhistle. Best to clarify what the person who typed it meant, before jumping to conclusions, sir. Not everyone is interested in filling their mind with extra rules just to cater to others' insecurities. Lastly, your comment violates the https://en.wikipedia.org/wiki/Principle_of_charity https://en.wikipedia.org/wiki/Principle_of_charity , which is a good principle for everyone to maintain.
- 4mo ago
- aliljet 4mo agoI'm curious about this. What models/tools have you been using?
- peterderivaz 4mo agoI've been trying out this model on a 4090 to transcribe a Japanese grammar pdf (written in English with lots of Japanese examples) and it seems to be working very well from the small parts I have double checked. The output contains both the kanji/hiragana and English as appropriate without attempting any translation. It has converted about 200 pages in an hour.