4 ms·
1. That's indeed dependant on the way the PDF has been created. Sometimes, you can even have letter splitted on different text token. 2. You're right, I mean M
by trez 13y ago
1. That's indeed dependant on the way the PDF has been created. Sometimes, you can even have letter splitted on different text token.
2. You're right, I mean Megabyte
- kijin 13y agoThanks for the clarifications. Since a lot of PDFs are badly organized (and I wonder if some programs deliberately do that to make text extraction difficult), perhaps you could try to analyze the location of each token on the page and merge the ones that seem to belong together. That would be already 100x better than most of the free PDF->text converters out there.
- alxbrun 13y agoOr even go the OCR approach !
- ygra 13y agoAll things considered that's pretty sad, though. A digital archive format that cannot reliably be read by machines, even if it contains just text.
- trez 13y agowe are already close to do that but with a really slow parser (this one can even replace some text on the pdf). Our problem now is to understand if developers would rather have better text extraction or some other features like image extractions, etc.. Let us know what you would prefer.
- kijin 13y agoImage extraction would be cool, but to me getting a readable block of text is more important.