3 ms·
Very cool. I think this is an interesting benchmarking task for a language model (as well as the practical uses). I tried the same thing some time ago, just on
by version_five 3y ago
Very cool. I think this is an interesting benchmarking task for a language model (as well as the practical uses). I tried the same thing some time ago, just on a random snippet of tesseract OCR. I had a Vicuna model (I forget which) that failed miserably, and chat GPT did it flawlessly. I did not have any hallucination problem with chatGPT.
It sounds from your writeup then like llama2 (which one) doesn't work well enough without some guardrails but it's possible to make it work? How would you rate the performance overall?
- eigenvalue 3y agoI’d say that it does work pretty well. It could simply be that I’m sampling too much from the LLM which is causing it to hallucinate more than it should, hence why I needed to spend so much time filtering out the hallucinations. But I think the risk of “made up” stuff in what’s supposed to be an accurate representation of a scanned document is big enough that you might always want to do something like that for quality control purposes, just to be on the safe side. I used the Llama2 13B Chat model ggml weights from TheBloke on Huggingface.
- version_five 3y agoYou should be able to build a text->image->corrupt with noise->tesseract pipeline to generate some synthetic supervised fine tuning examples (or just try corruption the text, I'm assuming putting it through tesseract makes more authentic corruption)
- eigenvalue 3y agoYeah that's a great idea. Would be a good excuse to learn about fine-tuning Llama2. Although corrupting with simple Gaussian noise might not be "hard" enough to help it learn how to recognize something like this: https://archive.org/details/royalnavyhistory01clow/page/402/mode/2up https://archive.org/details/royalnavyhistory01clow/page/402/... properly, which has specks/spots, parts that are much lighter/fainter, printing errors, etc.
- version_five 3y agoAs long as you can get tesseract to make similar text errors, the actual corruption doesn't matter though. You'd have to play with it but I'd guess the most important thing is to train the model to only guess the most likely words and avoid making anything up, as opposed to learning anything new about what words best fit a given corruption. So it may be enough even to try just randomly dropping letters from text and training with that. There's an interesting set of experiments there anyway.