4 ms·
The challenge I have is how to get bounding boxes for the OCR, for things like redaction/de-identification.
by techwizrd 2y ago
The challenge I have is how to get bounding boxes for the OCR, for things like redaction/de-identification.
- kbyatnal 2y agoyeah that's a fun challenge — what we've seen work well is a system that forces the LLM to generate citations for all extracted data, map that back to the original OCR content, and then generate bounding boxes that way. Tons of edge cases for sure that we've built a suite of heuristics for over time, but overall works really well.
- dontlikeyoueith 2y agoWhy would you do this and not use Textract?
- schcrosby 2y agoI too have this question.
- dontlikeyoueith 2y agoAWS Textract works pretty well for this and is much cheaper than running LLMs.
- daemonologist 2y agoTextract is more expensive than this (for your first 1M pages per month at least) and significantly more than something like Gemini Flash. I agree it works pretty well though - definitely better than any of the open source pure OCR solutions I've tried.
- yfontana 2y agoI'm working on a projet that uses PaddleOCR to get bounding boxes. It's far from perfect, but it's open source and good enough for our requirements. And it can mostly handle a 150 MB single-page PDF (don't ask) without completely keeling over.