4 ms·
I built something similar in the past using TesseractOCR and Apache Tika and PyPDF2 / QPDF. The idea is sound. An API based OCR already exists in Apple / Micros
by jamescampbell 5y ago
I built something similar in the past using TesseractOCR and Apache Tika and PyPDF2 / QPDF. The idea is sound. An API based OCR already exists in Apple / Microsoft / and Google so I am not sure this would be that useful. There would be no way for the user to trust that you are not taking the data you are OCR'ing and using it. If you can apply some type of one way encryption of the content and prove it via open source code (like Whisper Systems does for Signal) which seems like overkill and lots of effort for a free app.
- zmix 5y ago> An API based OCR already exists in Apple / Microsoft / and Google Where can I reach them? Thx.
- konfuzio 5y agohttps://centraluseuap.dev.cognitive.microsoft.com/docs/services/computer-vision-v3-2/operations/5d9869604be85dee480c8750 https://centraluseuap.dev.cognitive.microsoft.com/docs/servi... We have build a web client + REST API that allows to use the API for free for small personal projects. https://konfuzio.com/en/ocr-api/ https://konfuzio.com/en/ocr-api/ It supports handwriting, correction of HOCR text via the webrowser, automated language detection. We use the text to allow large enterprises to train document categorization and data extraction AI in a low/now code UI. Disclaimer: I'm one of the founders.
- zmix 5y agoThanks, I will check this out. Nice to explain the background also.