4 ms·
Made a small project to help extract structure from documents (pdf,jpg,etc -> JSON or CSV): https://datasqueeze.ai/ https://datasqueeze.ai/ There's 10 free pag
by Zaheer 2y ago
Made a small project to help extract structure from documents (pdf,jpg,etc -> JSON or CSV): https://datasqueeze.ai/ https://datasqueeze.ai/
There's 10 free pages to extract if anyone wants to give it a try. I've found that just sending a pdf to models doesn't extract it properly especially with longer documents. Have tried to incorporate all best practices into this tool. It's a pet project for now. Lmk if you find it helpful!
- artisandip7 2y agotried it works great, ty!
- matchagaucho 2y agoSimilarly I've found old-school OCR is needed for more reliability.
- bagels 2y agoCombining google's ocr with llm gives OCR superpowers. Tell the llm the text is from an ocr and ask it to correct it.
- saturn8601 2y agoThat sounds like it could be very dangerous when the LLM gets it wrong...
- bagels 2y agoDepends what you're using it for. If you're relying on OCR, you've already got to accept some amount of error.
- MarkMarine 2y agoI've been using this to OCR some photos I took of books and it's remarkable at it. My first pass was just a loop where I'd OCR, feed the text to the model and ask it to normalize into a schema but I found out just sending the image to the model and asking it to OCR and turn it into the shape of data I wanted was so much more accurate.
- deleted 2y ago[deleted]
- hackernewds 2y agoIs this simply the OCR bits to feed to openai structured output?