9 ms·
Sorry I was harsh. I tried the translation service and it work better. Anyway, I don't remember exactly all the articles I tried, but two of them were https:/
by pathsjs 11y ago
Sorry I was harsh. I tried the translation service and it work better.
Anyway, I don't remember exactly all the articles I tried, but two of them were
https://www.cs.princeton.edu/~chazelle/pubs/mst.pdf https://www.cs.princeton.edu/~chazelle/pubs/mst.pdf
https://www.cs.ubc.ca/~condon/papers/chungcondon96.pdf https://www.cs.ubc.ca/~condon/papers/chungcondon96.pdf
while one news that failed to parse was
http://www.repubblica.it/economia/2016/02/09/news/borse_9_febbraio-133008943/ http://www.repubblica.it/economia/2016/02/09/news/borse_9_fe...
I tried other articles and news, but I do not recall each of them exactly
- pesenti 11y agoWe will check them and get back to you. Thanks!
- mfulgo 11y agoHey pathsjs, sorry for the bad experience... Nevertheless, thanks for the feedback. TLDR: I pushed a fix for the bad character issue, and those PDFs should convert now. The long version: It has to do with the underlying structure of the PDF; some of the characters in the above PDF have glyphs for display but don't actually map the characters to code points. So, when we pull out the text, they come through as invalid characters, which we should have filtered out. This is an issue we've seen with (all?) PDF viewers; the text you copy from a sentence isn't always what you expect... But, we're aware of that shortcoming and are looking at some ways to improve the quality. In regard to the extra content in the news articles, we're not currently trying to do what BoilerPlate does. If you want to include or exclude specific content from a page, we have config options to do that via XPaths. Though, we're always open to ways of improving our services, and incorporating something like that would probably be useful.