4 ms·
tesseract is fine for basic use cases, but it fails when the image is tilted (and thus the text isn't laid out horizontally), which can happen several times wit
by felipefar 2y ago
tesseract is fine for basic use cases, but it fails when the image is tilted (and thus the text isn't laid out horizontally), which can happen several times with scanned books. Compared to how well the Google OCR engine works, tesseract should be much better than it is.
I wonder how difficult it is to develop a better OCR engine than tesseract.
- perihelions 2y agoAm I overlooking something, or is automating page rotation no more work than just a 2d FFT?
- notyoutube 2y agoMind ELI5ing this? it seems neat
- perihelions 2y agoThe Fourier transforms map plane waves to points. Blocks of regularly-spaced text have a periodic character, with the period length of their line spacing; their Fourier transform (I think??) would, in 2d frequency space, have amplitude peaks on vectors that have the same angle as the rotation of the lines.
- notyoutube 2y agothanks!
- aidenn0 2y agoI think it's more typical to low-pass (i.e. blur) the image and then use a line-detection algorithm like the Hough transform. Properly deskewed text should have prominent horizontal white lines.
- perihelions 2y ago- "Hough transform" Oh, that one has much nicer properties—thank you!
- HeatrayEnjoyer 2y agoTesseract is last gen. Multimodal is SOTA, and can handle even heavily distorted or destroyed text.
- aidenn0 2y agoYou are supposed to deskew (and de-warp if the image isn't flat) images before running through tesseract. There are other tools for doing that.