4 ms·
Hi, We used PdfTextStream for extracting information from pdf documents in a similar manner as you describe (pre-known regions of the document), after lookin
by atripathi 16y ago
Hi,
We used PdfTextStream for extracting information from pdf documents in a similar manner as you describe (pre-known regions of the document), after looking at few other options. It was not very easy though working with coordinates and rectangles though :)
We observed that the text in our pdf had a structure to it. So instead we simply dumped the text from pdf using pdftotext and wrote an ANTLR grammar for the structure we saw. This enabled us to parse relevant information from the text dump.