6 ms·
You should be able to build a text->image->corrupt with noise->tesseract pipeline to generate some synthetic supervised fine tuning examples (or just try corrup
by version_five 3y ago
You should be able to build a text->image->corrupt with noise->tesseract pipeline to generate some synthetic supervised fine tuning examples (or just try corruption the text, I'm assuming putting it through tesseract makes more authentic corruption)
- eigenvalue 3y agoYeah that's a great idea. Would be a good excuse to learn about fine-tuning Llama2. Although corrupting with simple Gaussian noise might not be "hard" enough to help it learn how to recognize something like this: https://archive.org/details/royalnavyhistory01clow/page/402/mode/2up https://archive.org/details/royalnavyhistory01clow/page/402/... properly, which has specks/spots, parts that are much lighter/fainter, printing errors, etc.
- version_five 3y agoAs long as you can get tesseract to make similar text errors, the actual corruption doesn't matter though. You'd have to play with it but I'd guess the most important thing is to train the model to only guess the most likely words and avoid making anything up, as opposed to learning anything new about what words best fit a given corruption. So it may be enough even to try just randomly dropping letters from text and training with that. There's an interesting set of experiments there anyway.