5 ms·
I was just playing with tesseract last week (I'd used it years ago) and wasn't too happy. I had a pretty simple pdf that was in what you could think of as an ol
by version_five 3y ago
I was just playing with tesseract last week (I'd used it years ago) and wasn't too happy. I had a pretty simple pdf that was in what you could think of as an old typewritten font, but easily legible, and I got all kinds of word fragments and nonsense characters in the output. I know that high quality ocr systems include a language model to coerce the read text into the most probable words. Is tesseract just supposed to be the first stage of such a system?
I'll note that when I put the tesseract output into chatgpt and prompted it saying it was ocr'd text and asking to clean it up, it worked very well.
- deleted 3y ago[deleted]
- flaviut 3y agoI was just processing a document with tesseract & ocrmypdf, and two things: My first time processing it, I used `ocrmypdf --redo-ocr` because it looked like there was some existing OCR. After processing, the OCR was crap because ocrmypdf didn't realize it was OCR but thought it was real text in the document that should be kept. This was fixable using `ocrmypdf --force-ocr`. Before realizing this, I discovered that Tesseract 4 & 5 use a neural network-based recognition. I then came across this step-by-step guide on fine-tuning Tesseract for a specific document set: https://www.statworx.com/en/content-hub/blog/fine-tuning-tesseract-ocr-for-german-invoices/ https://www.statworx.com/en/content-hub/blog/fine-tuning-tes... I didn't end up following the fine-tuning process because at this point `ocrmypdf --force-ocr` worked excellently, but I thought the draw_box_file_data.py script from their example was particularly useful: https://gist.github.com/flaviut/d901be509425098645e4ae527a9e9f3a https://gist.github.com/flaviut/d901be509425098645e4ae527a9e...
- denysvitali 3y agoFWIW, I'm using Google's ML Kit which runs completely on-device and doesn't send the documents to Google. It works better than tesseract for my use case. I did a presentation on the topic recently: https://clis-everywhere.k8s.best/16 https://clis-everywhere.k8s.best/16 I'll soon make the stack open source, but it shouldn't be hard to recreate given the inputs I've already provided.