4 ms·
As long as you can get tesseract to make similar text errors, the actual corruption doesn't matter though. You'd have to play with it but I'd guess the most imp
by version_five 3y ago
As long as you can get tesseract to make similar text errors, the actual corruption doesn't matter though. You'd have to play with it but I'd guess the most important thing is to train the model to only guess the most likely words and avoid making anything up, as opposed to learning anything new about what words best fit a given corruption. So it may be enough even to try just randomly dropping letters from text and training with that. There's an interesting set of experiments there anyway.