3 ms·
I'm doing this for a customer right now! Using Textract from Amazon to do the OCR. I dump the raw json in a SQLite database. Information extraction is trickie
by pani5ue 4y ago
I'm doing this for a customer right now!
Using Textract from Amazon to do the OCR. I dump the raw json in a SQLite database.
Information extraction is trickier. I extract useful things like lines and pages, locations of lines, set up various columns and now can search the SQLite db either with python code or SQL queries.
There's of course cost for running textract OCR on AWS. There are open source solutions like Tesseract
- chaps 4y agoOh yes, I'm very familiar with these. All these do though is extract information, but don't immediately make them useful. So there's a massive gulf of a middle-step that's not yet done. Textract gets close...ish to that, but it's prohibitively expensive.
- otoburb 4y agoEven with Amazon Textract, the middle step to curate extracted information into some form of meaning is still missing. Didn't realize this is still an unsolved problem.
- chaps 4y agoVery much so. Here's an example document I work with that has information that's difficult to extract from: https://s3.documentcloud.org/documents/6929951/CRID-1061543.pdf https://s3.documentcloud.org/documents/6929951/CRID-1061543.... Lots of missing context from these sheets that has to be interpreted (ie, how do you taxonomize each field of information?). Then asking questions on top of these documents is a step on top: "is the allegation about sexual violence?", "What is the name and rank of the person being accused?", "Is anything anomalous in the review process?", "Has this person's rank changed in the past 5 years?" etc etc. Now expand this problem to hundreds of thousands of different types of document.
- pani5ue 4y agoIf you haven't yet, you should look at the full text search capabilities of SQLite and postgres. Could simplify your search part of the workflow a bit