6 ms·
Claude 100k 1.3 blew me away. Giving it a task of extracting a specific column of information, using just the table header column text, from a table inside a P
by celestialcheese 3y ago
Claude 100k 1.3 blew me away.
Giving it a task of extracting a specific column of information, using just the table header column text, from a table inside a PDF, with text extracted using tesseract, no extra layers on top. (for those that haven't tried extracting tables with OCR, it's a non-trivial problem, and the output is a mess)
> 40k tokens in context, it performed at extracting the data, at 100% accuracy.
Changing the prompt to target a different column from the same table, worked perfectly as well. Changing a character in the table in the OCR context to test if it was somehow hallucinating, also accurately extracted the new data.
One of those "Jaw to the floor" moments for me.
Did the same task in GPT-4 (just limiting the context window to just 8k tokens), and it worked, but at ~4x more expensive, and without being able to feed it the whole document.
- anonymouse008 3y ago> text extracted using tesseract You're saying 'the text' without normalizing the rows and columns (basically the tab, space or newline delimited text with sporadic lines per row) was all you needed to send? I still have to normalize my tables even for GPT-4, I guess because I have weird merged rows and columns that attempt to do grouping info on top of the table data itself.
- celestialcheese 3y agoexactly. Just sent raw tesseract output, no formatting or "fix the OCR text" step. So the data looked like: ``` col1col2col3\nrow label\tdatapoint1\tdatapoint2... ``` Very messy. I don't think this is generalizable with the same 100% accuracy across any OCR output (they can be _really_ bad). I'm still planning on doing a first pass with a better Table OCR system like Textract, DocumentAI, PaddPaddle Table, etc which should improve accuracy.
- anonymouse008 3y agoThat’s still super cool! Yeah my use cases are in the really bad category - I’ve been building parsers for a while, and I’ve basically given up to manually stating rows of interest if present logic. Camelot got so close but I ended up building my own control layer to pdfminer.six to accommodate (I’d recommend Camelot if you’re still exploring). It absolutely sucks needing to be so specific out the gate, but at least the context rarely changes.
- pplante 3y agoWhat is the source of these nasty docs? I am also working on a layer above pdfminer.six to parse tables. It seems like this task is never done. LLMs have had mixed results for me too. I am focused on documents containing invoices, income statements, etc from the real estate industry. My email is in my profile if you want to reach out and compare notes!
- swyx 3y agobetter - you can do it copy pasting from pdf to gpt on your phone! https://twitter.com/swyx/status/1610247438958481408 https://twitter.com/swyx/status/1610247438958481408
- anonymouse008 3y agoDefinitely tried that way too, it didn’t work - my tables are pretty dang dumb. Merged cells, confidence intervals, weird characters in the cell field that change based on the row values - messing up a simple regex test, it’s really a billion dollar company solution but I’m about to punt it to the moon because it’s never fully done.
- arnaudsm 3y agoUsing LLMs with 100GB VRAM to convert PDFs to CSVs is truly depressing, but I am sure many companies will love it. 2023 office software already uses 1000x more ressources than 1990s'. I bet we are ready to do that again.
- martythemaniak 3y agoYou're missing the developer time. You no longer have to spend hours (or days, perhaps weeks depending on the sources) stringing together random libs, munging and cleaning data, testing, etc etc.
- arnaudsm 3y agoI agree, computers are cheapers than engineers. But I wonder how much more productive our economies could be if everyone was taught programming the same way we teach reading & writing, and open standards were ubiquitous.
- JumpCrisscross 3y ago> wonder how much more productive our economies could be if everyone was taught programming Prompt engineering is turning coding problems into language problems. It’s conceivable that humans writing code becomes artisanal in a century.
- vermilingua 3y agoCoding problems have always been language problems
- JumpCrisscross 3y ago> Coding problems have always been language problems Pedantically, sure. The field ChatGPT is most impactfully commoditizing is low-level coding. Instead of someone giving natural language instructions to a team of humans, they're increasingly able to give them to an LLM. It's an open question how far this can scale. But we may be near the zenith of the practicality of large-scale coding expertise.
- modernpink 3y agoWhat was the dollar cost to do this work? To iterate over a 40k context must be expensive.
- celestialcheese 3y ago~$0.45