4 ms·
I'd be interested in how people parse text off a PDF. I'm making a TTS tool to convert documents (mainly HTML docs at the moment) to speech, PDF's would be a gr
by radicalriddler 4y ago
I'd be interested in how people parse text off a PDF. I'm making a TTS tool to convert documents (mainly HTML docs at the moment) to speech, PDF's would be a great addition.
- airbreather 4y agouse pdftotext with the -bbox option to get the bounding box co-ords for each bit of text
- bradleykingz 4y agoHow I dealt with it was by first converting to HTML using PDF Box. Extracting text from there is then pretty simple.
- summarity 4y agoApache Tika can extract text (and metadata) from pretty much any file format ever invented.
- radicalriddler 4y agoThanks! I'll give this a shot. I wonder if it properly parses into real HTML or it's just div's all the way down.