3 ms·
Reading PDFs as html would Be nice, but as a ML engineer having high quality conversion would be very useful for large scale analysis and information extraction
by mendeza 7y ago
Reading PDFs as html would
Be nice, but as a ML engineer having high quality conversion would be very useful for large scale analysis and information extraction of pdf documents!
- ldenoue 7y agoWhat kind of analysis and extraction? Figures, tables? Other?
- udayrddy 7y agosame here. I mostly deal with text analytics, while the text PDFs do not create much issues, unless a crazy font is used, and the 2 column pages are a nightmare. In case you are looking for an API to extract structure rich content like tables from PDFs or images, look into this https://extracttable.com https://extracttable.com (p.s. I contributed to it)