10 ms·
Show HN: Extract Table from Image
- howmayiannoyyou 5y agoNice job. Actually though, what the world really needs in ML that divines the trend and perhaps indices/values from images of charts.
- plaidfuji 5y agoThis has been my pet side project for many years. What use case would you apply it to?
- howmayiannoyyou 5y agoScraping financial content
- MattGaiser 5y agoPair this with a snipping tool and all sorts of people in banking would use it for a few hours a day, especially if it could paste to Excel or at least fill the clipboard in a way pastable to Excel. I used to work for a bank on their innovation team and pitched basically this, but as an intern I had neither the skill nor time to do it. But it was certainly something a bunch of people internally wanted.
- deleted 5y ago[deleted]
- EveYoung 5y agoI can only imagine what a pain it would be to get InfoSec approval for such a tool, unless it's doing everything on-device.
- MattGaiser 5y agoWouldn’t need to be on device necessarily. At least my bank was comfortable with cloud everything and people using APIs from approved partners. If you can write the report in Google Docs, as long as they were the ones plugging in their API key for the OCR, I imagine it would be fine.
- saradhi 5y agoYou should consider extracttable.com P.s: I run the linked resource.
- v3gas 5y agoInteresting, thanks! Do you happen to know how to paste regular UTF-8 text into Excel/Google sheets as multiple cells? If I copy two cells in Sheets, I get a tab character (\t) between the cells. But if I try to paste "hello \t world" into Sheets then it's just dumped into one cell.
- v3gas 5y agoNevermind, the tab character is indeed what's needed to split it into multiple cells.
- jnsie 5y agoReally cool. I'm interested to hear your plans for this. Are you planning to offer as a service/open source/etc.?
- whirlwin 5y agoNice. Fun fact: The third example table is an ordered list of Norway's richest people (according to net worth, I think)
- w-m 5y agoI'm answering questions about Pandas (the Python data analysis framework) on StackOverflow from time to time. It's an exercise in patience, because many people will post screenshots of their data instead of a reproducible code example. You'll have to point about every other newcomer to the documentation on how write a proper question that one can actually answer. I'd imagine other areas around StackOverflow (SQL, R?) are fighting similar issues. I've just tried it with a question (sure enough the second newest Pandas tagged question had a table as an image), and your tool produced a nice .csv. It would be a godsend to have a button on StackOverflow that would replace a user-uploaded image of a table with some Pandas code that constructs the same DataFrame. Currently I would have to download the image, upload it to extract-table.com, download the .csv, load it into Python, run some code to create the code-based DataFrame. I'd consider sending people on StackOverflow to your tool if you cut down some of the steps: (1) allowing to paste in an URL of an image, and (2) producing Pandas code output that can be directly copy/pasted from the site (not having to download a csv). For illustration: here's what the Pandas code would look like for the first example of extract-table.com: df = pd.DataFrame( {'Name': {0: 'David', 1: 'Jessica', 2: 'Warren'}, 'Gender': {0: 'Male', 1: 'Female', 2: 'Male'}, 'Age': {0: 23, 1: 47, 2: 12}} )
- MattGaiser 5y agoCould do it with a Chrome extension. Add a button to the right click context menu and get the tabular data in the popup.
- pietrovismara 5y agoOff topic funny story: My highest voted answer on SO is a very basic one about Pandas, from 7 years ago. It's funny that I've only used Pands for a few weeks, years ago (I would need to relearn it from scratch now), but 90% of my SO score comes from that answer and I still get more points almost daily. In fact I'm in the top 6% of SO mostly thanks to that answer.
- belval 5y agoI'm in the same boat, 95% of my SO points come from an answer that was basically a copy pasted script to fix an obscure VMWare error with Ubuntu. Turns out a lot of people had the same issue that day.
- BillSaysThis 5y agoReally nice but... wondering how long this will last as a free tool given AWS fees.
- z3t4 5y agoShould make it into a browser plugin, so annoying when web sites have tables in images.
- mzs 5y agohttps://github.com/vegarsti/extract-table https://github.com/vegarsti/extract-table
- greaterweb 5y agoNice work putting together this tool. Have you seen either Spark OCR[1] from John Snow Labs or the Adobe PDF Extract API[2]? They both do a pretty good job a data extraction from tables as well. [1] https://www.johnsnowlabs.com/spark-ocr/ https://www.johnsnowlabs.com/spark-ocr/ [2] https://www.adobe.io/apis/documentcloud/dcsdk/pdf-extract.html https://www.adobe.io/apis/documentcloud/dcsdk/pdf-extract.ht...
- v3gas 5y agoThanks! No, I hadn't heard of either - thank you!
- pveierland 5y agoNeat tool! There appears to be two minor issues in the last example. There is an encoding issue of "ø" characters ("Røkke"), and a column split appears to be missing betweeen the closely spaced numbers ("33 300 22 700" vs "33 300,22 700"). Possible possibly non-trivial improvement: harmonize formatting within the same column to avoid mixed occurences of "7800" / "7 800".
- eihli 5y agoNice. I worked on something similar but far less robust: https://github.com/eihli/image-table-ocr https://github.com/eihli/image-table-ocr. It fails to find the tables on the example images at extract-table.com, but the code is heavily commented at https://eihli.github.io/image-table-ocr/pdf_table_extraction_and_ocr.html https://eihli.github.io/image-table-ocr/pdf_table_extraction... so there's high visibility into what's going on and what needs to change to get it to work with images of different sizes/fonts.
- nanis 5y agoWith this image[1] from this question on SO[2], the output[3] is missing the last row. FWIW, I've had the occasional miraculous-looking results from AWS Textract, but you do need to keep an eye on what's happening. Update: I just checked a bit carefully, and this example[4] is also missing the last row. Also, Danish ø seems problematic on your web page whereas the CSV has the right UTF-8 encoded bytes. [1]: https://i.stack.imgur.com/y7Zrt.png https://i.stack.imgur.com/y7Zrt.png [2]: https://stackoverflow.com/q/69363708/100754 https://stackoverflow.com/q/69363708/100754 [3]: https://results.extract-table.com/8d4818867ad604792819e98808ca447d2e1d33b3f69817a475a2d05c7a932e8e https://results.extract-table.com/8d4818867ad604792819e98808... [4]: https://results.extract-table.com/254d95722a2c2b1df72fc26b59925ef94d5c91017a661a194d76f1a52e228634 https://results.extract-table.com/254d95722a2c2b1df72fc26b59...
- v3gas 5y agoThat's interesting. Thanks for reporting!
- BrandiATMuhkuh 5y agoThis is really awesome. I have tried to solve that many times. I got close, with open CV and azure ML. I have even tried AWS Textract (~2 years ago). But this is the best implementation I have seen so far. Congratulations. I'm not sure what application you are thinking off. But the reason I'm following this problem is UX. Years ago, I worked on a project where anyone can add product prices into a DB. They do that by typing their receipt (line items) into the DB. The major issue was, the UX was horrible. With an API like yours, this is super simply. One photo. That's all. Maybe I'll revisit it as a side project.
- v3gas 5y agoThank you! I have also been kind of obsessed with this problem. I have tried to solve it myself, going from an image to bounding boxes and trying to separate the boxes into columns. But that problem is just fraught with edge cases, so I decided to just use an existing tool.
- tuberelay 5y agoUI Path does this in a nice way
- ducktective 5y agoAwesome project! Can AWS Textract be used directly with curl to return text strings of an uploaded image?
- v3gas 5y agoThanks! No, not that I know of, looks like for the AWS cli it needs to be in an S3 bucket, based on looking at this document: https://docs.aws.amazon.com/cli/latest/reference/textract/analyze-document.html https://docs.aws.amazon.com/cli/latest/reference/textract/an...
- ducktective 5y agohmm...weird. They could have provided a rate-limited API endpoint as a service...
- basmango 5y agoDoes it use textract directly? Or are you doing some preprocessing?
- v3gas 5y agoDirectly, no preprocessing! The postprocessing is concatenating all words that belong to the same cell.
- visarga 5y agoDoes it also do table detection in a larger image and header/body classification?
- v3gas 5y agoThis currently returns an error if it doesn't find exactly one table in the image, so it might be able to work with larger images, but probably not if there are multiple distinct blocks of text.