5 ms·
I wonder how LlamaParse compares head to head with https://unstructured.io https://unstructured.io
by johnsutor 3y ago
I wonder how LlamaParse compares head to head with https://unstructured.io https://unstructured.io
- justanotheratom 3y agonot clear to me why this got downvoted. sensible question.
- infecto 3y agoI would also like how it compares to any of the commercial offerings from Azure/AWS/GCP. They all have document parsing tools that I have found better than tools like unstructured. Sure you don't have some of the "magic" of segmenting text for vectorization and RAG but imo thats the easy part. The hard part is pulling data, forms, tables, text out of the PDF which I find the cloud tools to do a superior job.
- bugglebeetle 3y agoNothing beats Google text extraction tasks in my testing, especially for East Asian languages. I wish something else worked better, because their services are fairly expensive.
- infecto 3y agoI mostly have used Textract but have found all 3 to be fairly similar in accuracy with differences in structure and compatible languages. I think my call out here is that I don't think any of these libraries, llamaindex or unstructured, can compete in this area. I would rather use GCP/Azure/AWS to define the structure from a PDF and these for the rag portion if anything.
- youngNed 3y agoBe careful with unstructured: https://github.com/Unstructured-IO/unstructured/blob/d11c70cf83fdb8a08fed2cf01c6c0bd114d817df/unstructured/utils.py#L287-L319 https://github.com/Unstructured-IO/unstructured/blob/d11c70c... from: https://github.com/open-webui/open-webui/issues/687 https://github.com/open-webui/open-webui/issues/687