10 ms·
Show HN: PDF to MD by LLMs – Extract Text/Tables/Image Descriptives by GPT4o
I've developed a Python API service that uses GPT-4o for OCR on PDFs. It features parallel processing and batch handling for improved performance. Not only does it convert PDF to markdown, but it also describes the images within the PDF using captions like `[Image: This picture shows 4 people waving]`.
In testing with NASA's Apollo 17 flight documents, it successfully converted complex, multi-oriented pages into well-structured Markdown.
The project is open-source and available on GitHub. Feedback is welcome.
- jdross 2y agoHow does this compare with commercial OCR APIs on a cost per page?
- yigitkonur35 2y agoIt is a lot cheaper! While cost-effectiveness may not be the primary advantage, this solution offers superior accuracy and consistency. Key benefits include precise table generation and output in easily editable markdown format. Let's make some numbers game: - Average token usage per image: ~1200 - Total tokens per page (including prompt): ~1500 - [GPT4o] Input token cost: $5 per million tokens - [GPT4o] Output token cost: $15 per million tokens For 1000 documents: - Estimated total cost: $15 This represents excellent value considering the consistency and flexibility provided. For further cost optimization, consider: 1. Utilizing GPT4 mini: Reduces cost to approximately $8 per 1000 documents 2. Implementing batch API: Further reduces cost to around $4 per 1000 documents I think it offers an optimal balance of affordability & reliability. PS: One of the most affordable solution on market, cloudconvert charges ~30$ for 1K document (pdftron mode required 4 credits)
- johndough 2y ago> I think it offers an optimal balance of affordability & reliability. It is hard to trust "you" when ChatGPT wrote that text. You never know which part of the answer is genuine and which part was made up by ChatGPT. To actually answer that question: Pricing varies quite a bit depending on what exactly you want to do with a document. Text detection generally costs $1.5 per 1k pages: https://cloud.google.com/vision/pricing https://cloud.google.com/vision/pricing https://aws.amazon.com/textract/pricing/ https://aws.amazon.com/textract/pricing/ https://azure.microsoft.com/en-us/pricing/details/ai-document-intelligence/ https://azure.microsoft.com/en-us/pricing/details/ai-documen...
- yigitkonur35 2y agoYou've got a point, but try testing it on a tricky example like the Apollo 17 document - you know, with those sideways tables and old-school writing. You'll see all three non-AI services totally bomb. Now, if you tweak it to batch = 1 instead of 10, you'll notice there's hardly any made-up stuff. When you dial down the temperature close to zero, it's super unlikely to see hallucinations with limited context. At worst, you might get some skipped bits, but that's not a dealbreaker for folks looking to feed PDFs into AI systems. Let's face it, regular OCR already messes up so much that...
- Propelloni 2y ago> you might get some skipped bits, but that's not a dealbreaker for folks looking to feed PDFs into AI systems Unless it is. We have a few hundred PDF per month (mostly tables) where we need 100% accuracy. Currently we feed them into an OCR and have humans check the result. I do not win anything if I have to check the LLM output, too.
- llm_trw 2y agoI'm currently solving this problem for work and thinking of a spin out, what's a ballpark figure you'd be willing to pay per 1000 pages for 99.999% character level accuracy?
- Lerc 2y agoI guess it depends on the use case, but if it surpasses the error rate that exists in the source document then it would be difficult to argue against. Specific things like evidentiary use would want 100% but that's at a level where any document processing would be suspect. What is the the typical range for error rate in PDF generation in various fields? Even robust technical documents have the occasional typo.
- llm_trw 2y agoI'm not using generative models to fill in details not present in the original document. If there's a typo there then there will be a typo in the transcript. If you want to fix that then you can run another model on top of it.
- smusamashah 2y agoI have not found any mention of accuracy. Since it's using LLM, how accurate the conversion is? As in does that NASA document match 100% with the pdf or did it introduce any made up things (hallucinations)? That converted NASA doc should be included in repo and linked in readme if you haven't already.
- yigitkonur35 2y agoPeople are really freaked out about hallucinations, but you can totally tackle that with solid prompts. The one in the repo right now is doing a pretty good job. Keep in mind though, this project is all about maxing out context for LLMs in products that need PDF input. We're not talking about some hardcore archiving system for the Library of Congress here. The goal is to boost consistency whenever you're feeding PDF context into an LLM-powered tool. Appreciate the feedback, I'll be sure to add that in.
- deleted 2y ago[deleted]
- williamdclt 2y agoI don’t think any prompting skill guarantees the absence of hallucination. And if hallucination is possible, you will usually need to worry about it
- Foobar8568 2y agoAs soon as you have something else than a paragraph in a single column layout, you will get hallucinations, random stuff, cut off etc even if you say which pages to look at, LLM will just do what ever.
- freedomben 2y agoCan you give some examples of prompts that you use that will tackle hallucinations?
- AdieuToLogic 2y ago> People are really freaked out about hallucinations, but you can totally tackle that with solid prompts. > The goal is to boost consistency whenever you're feeding PDF context into an LLM-powered tool. These two assertions are contradictory. There are no "solid prompts" which obviate anthropomorphic "LLM hallucinations." Also, there is no deterministic consistency when "feeding PDF context" into an intrinsically non-deterministic algorithm, as any "LLM-powered tool" is by definition.
- magicalhippo 2y agoWas just looking for something like this. Does it handle equations to latex or similar? How about rotated tables, ie landscape mode but page is still portait?
- yigitkonur35 2y agoI messed around with some rotating tables in that Apollo 17 demo video - you can check it out in the repo if you want. It's pretty straightforward to tweak just by changing the prompt. You can customize that prompt section in the code to fit whatever you need. Oh, and if you throw in a line about LaTeX, it'll make things even more consistent. Just add it to that markdown definition part I set up. Honestly, it'll probably work pretty well as is - should be way better than those clunky old OCR systems.
- nicodjimenez 2y agoHave you checked out Mathpix? It's another option. Disclaimer: I'm the founder.
- troysk 2y ago+1! Most LLMs can already output Mathpix markdown. I prompt it to do so and it gives the code and this use a rendering library to show the scalable and selectable equations. No wonder facebook nougat also uses it. Good stuff!
- magicalhippo 2y agoWas looking for a self-hosted solution as I have quite on/off needs, but I'll give it a whirl as it looks quite promising.
- x_may 2y agoCheck out Nougat from meta
- magicalhippo 2y agoThanks, looks very interesting, but also somewhat abandoned. Will keep an eye on it in case someone picks up the torch.
- jdthedisciple 2y agoZerox [0] was featured on here recently and does the exact same thing [0] https://github.com/getomni-ai/zerox https://github.com/getomni-ai/zerox
- pierre 2y agoParsing docs using LVM is the way forward (also see OCR2 paper released last week, people are having ablot of success parsing with fine tunned Qwen2). The hard part is to prevent the model ignoring some part of the page and halucinations (see some of the gpt4o sample here like the xanax notice:https://www.llamaindex.ai/blog/introducing-llamaparse-premium https://www.llamaindex.ai/blog/introducing-llamaparse-premiu...) However this model will get better and we may soon have a good pdf to md model.
- authorfly 2y agoWhat about combining old school OCR with GPT visual OCR? If your old school OCR output has output that is not present in the visual one, but is coherent (e.g. english sentences), you could get it back and slot it into the missing place from the visual output.
- yigitkonur35 2y agoYou're absolutely right. I use PDFTron (through CloudCovert) for full document OCR, but for pages with fewer than 100 characters, I switch to this API. It's a great combo – I get the solid OCR performance of SolidDocument for most content, but I can also handle tricky stuff like stats, old-fashioned text, or handwriting that regular OCR struggles with. That's why I added page numbers upfront.
- fkilaiwi 2y agowhat paper are you referring to?
- perrywky 2y agoI guess this: https://arxiv.org/html/2409.01704v1 https://arxiv.org/html/2409.01704v1
- fzysingularity 2y agoWe’ve been doing exactly this by doubling-down on VLMs (https://vlm.run https://vlm.run) - VLMs are way better at handling layout and context where OCR systems fail miserably - VLMs read documents like humans do, which makes dealing with special layouts like bullets, tables, charts, footnotes much more tractable with a singular approach rather than have to special case a whole bunch of OCR + post-processing - VLMs are definitely more expensive, but can be specialized and distilled for accurate and cost effective inference In general, I think vision + LLMs can be trained to explicitly to “extract” information and avoid reasoning/hallucinating about the text. The reasoning can be another module altogether.
- bravura 2y agoI've also been using the nougat models from meta, which are trained to turn PDF into md using the donut architecture
- rahimnathwani 2y agoDoes it work well on documents that aren't academic papers? https://facebookresearch.github.io/nougat/ https://facebookresearch.github.io/nougat/
- zerop 2y agoI had been using GPT4o for extracting insights from Scanned docs, it was doing fine. But very recently (since they launched new model - o1), it's not working. GPT4o is refusing to extract text from images and says it can't do it, though it was doing same thing with same prompts till last week. I am not sure if this is intentional downgrade and it can be clubbed with new model launch, but it's really frustrating for me. I cancelled my GPT4 premium and moved to claude. It works good.
- fragmede 2y agoWhy not just switch back to GPT-4? it's still there.
- itchyjunk 2y agoSo is 4o. Problem isn't the absence of model, it's inconsistency.
- itissid 2y agoThis. Inconsistency is a big problem for large tasks, you are better off making your own models to do this. I have seen this odd kind of inconsistency in generating the same results, sometimes in the same chat itself after starting off fine. I was once trying to extract hand written dates and times from a large pdf document in batches of 10 pages at a time from a very specific part of the page. IN some documents it started by refusing, but not in other different chat windows that I tried with the same document. Sometimes it would say there is an error, and then it would work in a new chat window. But I am not sure why, but just starting a new chat works for these kind of situations. Sometimes it will start off fine with OCR, then as the task progresses, it will start hallucinating. Even though the text to be extracted follows a pattern like dates, it for the life of me could not get it right.
- rrrix1 2y ago> "...you are better off making your own models to do this" I'm doubtful you meant what you wrote here. Using a readymade UI or API to perform an effectively magical task (for most of us) is an entirely different paradigm to "just train your own model." In reality, for us non-ML model training mortals, we're actually probably better off hiring a human to do basic data entry.
- eth0up 2y agoI used GPT4o to convert heavily convoluted PDFs into csv files. The files were Florida Lottery Pick(n) histories, which they deliberately convolute to prevent automatic searching; ctrl-f does nothing and a fsck-ton of special characters embellish the whole file. I had previously done so manually, with regex, and was surprised with the quality of the end results of GPT, despite many preceding failed iterations. The work was done in two steps, first with pdf2text, then python. I'm still trying to created a script to extract the latest numbers from the FL website and append to a cvs list, without re-running the stripping script on the whole PDF every time. Why? I want people to have the ability to freely search the entire history of winning numbers, which in their web hosted search function, is limited to only two of 30+ years. I know there's a more efficient method, but I don't know more than that.
- alchemist1e9 2y agoOff topic - but the obvious follow up question is why do you want people to have this ability to search the entire history?
- eth0up 2y agoThanks for asking... 1) I'm a rebel 2) I am irritated by deliberate obfuscations of public data, especially by a source that I suspect is corrupt. Although my extensive analysis has not yet revealed any significant pattern anomalies in their numbers. 3) It's kind of my re-intro into python, which I never made significant progress in but always wanted to. 4) It's literally the real history of all winning numbers since inception. Individuals may have various reasons for accessing this data, but I've been using it to test for manipulation. I presume for most folks it would be curiosity, or gambler's fallacy type stuff. Regardless, it shouldn't be obfuscated.
- alchemist1e9 2y agoI had suspected you’re are suspicious of manipulation. I have heard many rumors of lottery corruption and manipulation. It’s certainly a big red flag if they are deliberately obstructing access to the data. Make sense your project and I’d probably take 30 mins to look at the data if I came across it. I’m somewhat decent at data and number analysis so if there is something and enough people can easily take a look at it, then it might get exposed. Interesting and good luck.
- scottmcdot 2y agoDoes it do image to MD too?
- deleted 2y ago[deleted]
- yigitkonur35 2y agoYes, you can customize this as you wish by adding it to your prompt.
- vunderba 2y agoUnless the only thing you want is a description of the image, then the real answer is NO. You can get an LLM to do something like "If you encounter an image that is not easily convertable to standard markdown, insert a [[DESCRIPTION OF IMAGE]] here." placeholder, but at that point you've lost information that may be salient to the original PDF. The reason is because these multimodal LLMs can give you descriptions/OCR/etc., but they cannot give you quantifiable information related to placement. So if there was a picture of a tiger in the middle of the page converted to a bitmap, you couldn't get the LLM to give you something like this: "Image detected at pixel position (120, 200) - (240, 500)." - because that's really want you want. You almost need segmentation system middleware that the LLM can forward to which can cut out these images to use in markdown syntax: 
- gdevenyi 2y ago[flagged]
- deleted 2y ago[deleted]
- charlie0 2y agoI do this all the time for old specs, but one issue I worry about is accuracy. It's hard to confirm if the translations are 100% correct.
- Oras 2y agoWhile this is a nice development, it’s quite risky parsing documents with LLMs. In usual OCRs, you have boundaries to check, but with LLMs, you just get a black box output. As others mentioned, consistency is key in parsing documents and consistency is not a feature of LLMs. The output might look plausible, but without proper validation this is just a nice local playground that can’t make it to production.
- yigitkonur35 2y agoI get your worries about LLMs and their consistency problems. But I think we can fix a lot of that using LLMs themselves for checks. If you're after top-notch accuracy, you could throw in another prompt, add some visual and text input, and double-check that nothing's lost in translation. The cheaper models are actually great for this kind of quality control. LLMs have come a long way since they first showed up, and I reckon they've stepped up their game enough to shake off that bad rap for giving mixed signals.
- Oras 2y agoHow would you know something is missing? I tried multiple OCRs before and it’s hard to tell if the output is accurate or not but just comparing manually. I created a tool to visualise the output of OCR [0] to see what’s missing and there are many cases that would be quite concerning especially when working with financial data. This tool wouldn’t work with LLMs as they don’t return the character recognition (to my knowledge), which will make it harder to evaluate them on a scale. If I want to use LLMs for the task, I would use them to help with training ML model to do OCR better, such as creating thousands of synthetic data to train. [0] https://github.com/orasik/parsevision https://github.com/orasik/parsevision
- yigitkonur35 2y agoWow, you knocked it out of the park! I'll be sure to use this when I tackle that evaluation.
- whiplash451 2y ago
- refulgentis 2y agoGPT 4o doesn't do actual OCR and there's much smaller and more effective models for specifically this problem. I appreciate your work, intent, and sharing it. It's very important to appreciate what you're doing and its context when sharing it. At that point, you are responsible for it, and the choices you make when communicating about it reflect on you.
- yigitkonur35 2y agoI've found this method really useful for prepping PDFs before running them through AI. I mix it with traditional OCR for a hybrid approach. It's a game-changer for getting info from tricky pages. Sure, you wouldn't bet the farm on it for a big, official project, but it's still pretty solid. If you're willing to spend a bit more, you can use extra prompts to check for any context skips. It's a lot of work, though - probably best left to companies that specialize in this stuff. I've been testing it out on pitch decks made in Figma and saved as JPGs. Surprisingly, the LLM OCR outperformed top dogs like SolidDocuments and PDFtron. Since I'm mainly after getting good context for the LLM from PDFs, I've been using this hybrid setup, bringing in the LLM OCR for pages that need it. In my book, this API is perfect for these kinds of situations.
- TZubiri 2y agoOk attempt number 158 at parsing pdfs, here we go, this time surely it will work.
- wittjeff 2y agolicense please?
- fzysingularity 2y agoOne nit in the repo README - you might want to change the cost reporting to be as $15 / 1000 pages instead of documents.
- constantinum 2y agoThere is also LLMWhisperer, a document pre-processor specifically made for LLM consumption. As other mentioned, accuracy is the one part of solution criteria, other include, how does the preprocessing engine scale/performs at large scale, and how does it handle very complex documents like, bank loan forms with checkboxes, IRS tax forms with multi-layered nested tables etc. https://unstract.com/llmwhisperer/ https://unstract.com/llmwhisperer/ LLMWhisperer is a part of Unstract - An open-source tool for unstructured document ETL. https://github.com/Zipstack/unstract https://github.com/Zipstack/unstract
- devops000 2y agoYou could transform arXiv to a markdown website
- KoolKat23 2y agoThis is handy, one thing I've noticed using 3.5 Sonnet, the tables that aren't the correct orientation are more prone to incorrect output. I know this was an issue when GPT 4 vision initially came out due to training, not sure if it's a solved problem or if your tool handles this.
- bschmidt1 2y agoMy previous employer needs this. I won't tell them :) :D >:D :|