5 ms·
It is a lot cheaper! While cost-effectiveness may not be the primary advantage, this solution offers superior accuracy and consistency. Key benefits include pre
by yigitkonur35 2y ago
It is a lot cheaper! While cost-effectiveness may not be the primary advantage, this solution offers superior accuracy and consistency. Key benefits include precise table generation and output in easily editable markdown format.
Let's make some numbers game:
- Average token usage per image: ~1200
- Total tokens per page (including prompt): ~1500
- [GPT4o] Input token cost: $5 per million tokens
- [GPT4o] Output token cost: $15 per million tokens
For 1000 documents:
- Estimated total cost: $15
This represents excellent value considering the consistency and flexibility provided. For further cost optimization, consider:
1. Utilizing GPT4 mini: Reduces cost to approximately $8 per 1000 documents
2. Implementing batch API: Further reduces cost to around $4 per 1000 documents
I think it offers an optimal balance of affordability & reliability.
PS: One of the most affordable solution on market, cloudconvert charges ~30$ for 1K document (pdftron mode required 4 credits)
- johndough 2y ago> I think it offers an optimal balance of affordability & reliability. It is hard to trust "you" when ChatGPT wrote that text. You never know which part of the answer is genuine and which part was made up by ChatGPT. To actually answer that question: Pricing varies quite a bit depending on what exactly you want to do with a document. Text detection generally costs $1.5 per 1k pages: https://cloud.google.com/vision/pricing https://cloud.google.com/vision/pricing https://aws.amazon.com/textract/pricing/ https://aws.amazon.com/textract/pricing/ https://azure.microsoft.com/en-us/pricing/details/ai-document-intelligence/ https://azure.microsoft.com/en-us/pricing/details/ai-documen...
- yigitkonur35 2y agoYou've got a point, but try testing it on a tricky example like the Apollo 17 document - you know, with those sideways tables and old-school writing. You'll see all three non-AI services totally bomb. Now, if you tweak it to batch = 1 instead of 10, you'll notice there's hardly any made-up stuff. When you dial down the temperature close to zero, it's super unlikely to see hallucinations with limited context. At worst, you might get some skipped bits, but that's not a dealbreaker for folks looking to feed PDFs into AI systems. Let's face it, regular OCR already messes up so much that...
- Propelloni 2y ago> you might get some skipped bits, but that's not a dealbreaker for folks looking to feed PDFs into AI systems Unless it is. We have a few hundred PDF per month (mostly tables) where we need 100% accuracy. Currently we feed them into an OCR and have humans check the result. I do not win anything if I have to check the LLM output, too.
- llm_trw 2y agoI'm currently solving this problem for work and thinking of a spin out, what's a ballpark figure you'd be willing to pay per 1000 pages for 99.999% character level accuracy?
- Lerc 2y agoI guess it depends on the use case, but if it surpasses the error rate that exists in the source document then it would be difficult to argue against. Specific things like evidentiary use would want 100% but that's at a level where any document processing would be suspect. What is the the typical range for error rate in PDF generation in various fields? Even robust technical documents have the occasional typo.
- llm_trw 2y agoI'm not using generative models to fill in details not present in the original document. If there's a typo there then there will be a typo in the transcript. If you want to fix that then you can run another model on top of it.
- Lerc 2y agoI realise that. The point is that a user is implicitly committing to the baseline error rate that exists in whatever means by which the document was created. If any additional loss was insignificant in proportion to that error rate then it would be unreasonable to reject it on that basis.