3 ms·
I've not found one either. I did this at a very large scale recently and ended up just using pdfplumber. I did POC Table Transformer but the cost was too high a
by cle 2y ago
I've not found one either. I did this at a very large scale recently and ended up just using pdfplumber. I did POC Table Transformer but the cost was too high at my scale--there are probably better options now anyway. Most seem to focus on structure detection and then use traditional OCR for the actual content extraction.
It's a very hard space in the long tail, like tables that span pages or tables with complex internal structures. I went into it thinking "eh how hard can tables be?". Very hard. Thankfully it's a pretty active research area.