4 ms·
I created a PDF table extractor tool last year with the same idea that it should be local only. Try it here: https://pdftableutil.possiblenull.com/app/ https://
by mgm__ 6y ago
I created a PDF table extractor tool last year with the same idea that it should be local only. Try it here: https://pdftableutil.possiblenull.com/app/ https://pdftableutil.possiblenull.com/app/
Also as a Google Docs addon (still local only) https://workspace.google.com/marketplace/app/pdf_table_importer/646940040599 https://workspace.google.com/marketplace/app/pdf_table_impor...
I had a bad case of scope creep, so the tool can also extract tables from scanned/image PDFs using OpenCV.js and tesseract OCR wasm build!
- redman25 6y agoThis is interesting. How accurate would you say it is?
- mgm__ 6y agoI haven't seen anything better. It started as a PoC and I decided not to include table detection on the page and require the user to draw box around the table. I use Tabula under the hood for the cell/row detection and it is really good given the correct mode is selected for the type of table. The modes are stream (find cells by spacing) or lattice (find cells by ruling lines). The OCR/OpenCV seemed to be fine as well as long as the text isn't too blurry. Here is a GIF of the OCR/OpenCV running on an example Image PDF: https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/AAAAAAAAEAg/q__NHhsbhb0WLl1KYuSPBt16b0OyY3stwCLcBGAsYHQ/s640-w640-h400/Kvgkz4NcqO.gif https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/A...
- kickbeak 6y agoWow That looks awesome, what did you use to display the PDF in the Browser? feels all really responsive!