3 ms·
If you've ever extracted json data from files using llms, you know the pain in deciding if you can trust it or not. We found a solution that works well for our
by TimurKramar 20d ago
If you've ever extracted json data from files using llms, you know the pain in deciding if you can trust it or not. We found a solution that works well for our users: we track the full lineage of every extracted filed from the source document to the structured output. Our approach takes into consideration ocr (parsing), extraction and transformations steps.
Results come with confidences attached, so you know which fields need a double check and approval from a human. And this human-in-the-loop mechanism, where humans only check low confidence results,. can be used effectively to speed up existing backoffice operations that deal with document workflows.
You can try it right now at openparser.dev , or read this (https://www.openparser.dev/blog/how-openparser-uses-lineage https://www.openparser.dev/blog/how-openparser-uses-lineage) on how we built it and why one-shotting an llm to extract data from documents simply didn't work for us and our users.
oh and if you want to build with full lineage yourself, we opensourced it: https://github.com/eigenpal/openparser/tree/main/packages/lineage https://github.com/eigenpal/openparser/tree/main/packages/li...