3 ms·
We would need more context/information about your specific objectives. - document conversion (pdftotext, pdfbox, apache tabula, etc.) - OCR (tesseract, pypdfo
by ahljoh 10y ago
We would need more context/information about your specific objectives.
- document conversion (pdftotext, pdfbox, apache tabula, etc.)
- OCR (tesseract, pypdfocr, etc.)
- Named-Entity-Recognition (NER) i.e. finding and recognizing entities in text (DBPedia Spotlight, stanford NER via NLTK, spacy)
- coreference resolution, dependency parsing (spacy, syntaxnet)
- abc03 10y agoThanks. Some great keywords to investigate. I'm namely interested in two areas at the moment: - invoices (I guess NER would be partially an Option) - web scrapping (wrapper induction)