5 ms·
In-Browser OCR
- amelius 5y agoTesseract-based.
- no_time 5y agoResults for even the most plain and recognizable cases (english language screenshots) are absolutely terrible. This aligns with my other experiences with Tesseract. In fact the only OCR I ever had any success with is ABBYY.
- amelius 5y agoYeah I tried Tesseract with screenshots from the Spotify UI, but with terrible results. Makes you wonder where the competition is. Is anyone building an OCR based on DL?
- tkgally 5y agoIf you upload an image file to Google Drive and right-click on it to open the file in Google Docs, you get very accurate OCR of the text in the image. I’ve used it a lot for both English and Japanese, and it has consistently worked very well. I just tried it with the not-great-quality image at https://www.geograph.org.uk/photo/1587043 https://www.geograph.org.uk/photo/1587043 . The result for the first four lines was: “This building was erected around 1874 to provide a location for a seismometer of the British Association. A seismometer is an instrument which is designed to record earthquarkes and the one located in this building was only one of a series of such instruments located in the vicinity of Comnie to investigate the earthquakes which had been, and continue to be, prevalent in the area.” The only mistake seems to be “Comnie” instead of “Comrie.” (The misspelling “earthquarkes” appears in the original.) In contrast, the In-Browser OCR gives the following for the same four lines: “m .me W m m: swm mm m mm .. lucauan m a smmmm. m m. anus» Lssm'm'm» A swmmm ‘5 3,. WWW mm. vs damned .a mom “mum: m m: an» army! m a... “mm was Only an. a: .1 gem a; 5m m<llumems mm m we mm 0. Camus w WNW: we earmquakes Wm»... mm W. m coulmu: m be Wax/Nam m m M”
- amelius 5y agoInteresting. Could it recognize the "EARTHQUAKE HOUSE" title?
- tkgally 5y agoYes. Despite the unusual font, Google rendered it perfectly.
- amelius 5y agoOk, I tried your benchmark in Tesseract 4.0.0-beta.1 in Ubuntu 18.04, and it gave me: > EARTHQUAKE HOUSE > ‘his building was erected around 1874 to provide a location for a seismometer of the British Association. A seismometer is an instrument which is designed to record earthquarkes and the ‘one located in this building was only one of a series of such instruments located in the vicinity ‘earthquakes which had been, and continue to be, prevalent in the area. > of Comrie to investigate Clearly, Tesseract 4.0 has problems following the baseline of the text. But otherwise, it is much better than the output from the website and even got the title correct. Which makes me think they use an older version (?) My commandline: tesseract 1587043_dcd093c4.jpg output -l eng
- slacka 5y agoThis project is badly in need of an update. It's using an ancient 4 yo. version of tesseract.js. The current version: https://github.com/naptha/tesseract.js https://github.com/naptha/tesseract.js is based on tesseract v4.1.1, a newer your Ubuntu 18.04's. The 4.0 version added new neural network system based on LSTMs, with major accuracy gains. https://fossies.org/diffs/tesseract/4.0.0_vs_4.1.0/ChangeLog-diff.html https://fossies.org/diffs/tesseract/4.0.0_vs_4.1.0/ChangeLog...
- amelius 5y agoOk, I tried the demo page of project naptha on exactly the same image, and it gave me: > ARTHQUAKE HOUSE 1 > Ths builing was erecied Ground 1874 to provide a location for a seismomater of the Bitish Aasociation. A sefémomatar is an nsirument which is designed 10 racord earthquarkes and the ne ocate n his buiding was only one of a seies of such instruments located in the vicinity of Coms tonvestgete the earthauakes which had been, and continue to be, prevalet n the area. The 4.0.0 version looked better to me ...
- kranner 5y agoTesseract 4.0 onwards uses an engine based on LSTM networks.
- patentatt 5y agoMy impression is that the open source tesseract could be a component of an OCR system, but much more has to be done in preprocessing, registration, and segmentation to be able to use it for what we could consider “OCR”. Quite disappointing that OCR is still something that is proprietary, closed source, and expensive in 2021.
- jonatron 5y agoTwo alternatives, which are designed for OCR from photos: https://github.com/PaddlePaddle/PaddleOCR/ https://github.com/PaddlePaddle/PaddleOCR/ https://github.com/JaidedAI/EasyOCR/ https://github.com/JaidedAI/EasyOCR/ It's worth trying them if Tesseract isn't giving you good accuracy.
- eastendguy 5y agoABBY is still very good, but for some languages other tools are meanwhile better or at least much cheaper/free: Google Cloud OCR, Amazon Textract, Azure OCR, OCR.space
- no_time 5y agoI use ABBYY specifically because I'd rather not give data/money to FAANG.
- jitl 5y agoProject Naptha [https://projectnaptha.com/ https://projectnaptha.com/] delivers a much more impressive result, but isn’t open source. Naptha uses text detection and extraction to isolate text from the image, which greatly improves accuracy.
- chromatin 5y agoWhat is the top of the line, publicity available (open source) OCR engine now available? Tessaract still? Not interested in proprietary cloud solutions.
- ChemSpider 5y agoYes. It is the very best open source engine. Which is easy, since it is the only one ;)
- 2Gkashmiri 5y agoAnyone knows a PDF OCR tool? I am using a free online one. I take pictures with opennotescanner which spits out a PDF which I want searchable. Tesseracts expects PNG and outputs to text. I want the same PDF to be hidden overlaid with OCR text. This free PDF online service does a decent job but offline would be better.
- deleted 5y ago[deleted]
- eastendguy 5y agoFree, but also online: https://ocr.space/searchablepdf https://ocr.space/searchablepdf It comes with an API, so you can integrate it with your workflow.
- 2Gkashmiri 5y agoI tried that. Their free api is too little of use. 1MB file. I use PDF24.org which actually does everything great. Anything offline?
- eastendguy 5y agoFree & offline: Tesseract + pdfsandwich, see http://www.tobias-elze.de/pdfsandwich/ http://www.tobias-elze.de/pdfsandwich/