3 ms·
No need to go on, that's already more than I bargained for, thank you so much for taking the time to respond :) I suspect we're using different PDF viewers; I'
by marvy 6y ago
No need to go on, that's already more than I bargained for, thank you so much for taking the time to respond :)
I suspect we're using different PDF viewers; I'm using the one that comes with my browser.
The results are somewhat less disastrous than you describe, but still bad. (I'm seeing maybe half the problems you mentioned.) I'm a bit curious which viewer you're using.
I can describe what happens when I copy/paste the stuff you mentioned on pages 2/8/etc., and if you're interested I will, but if you're not interested let me ask a slightly different question:
Rather than try to get the screen reader to make sense of the final PDF, would it be easier to just download the original page source from arxiv and let the screen reader deal with that?
- j-pb 6y agoI was using the builtin viewer of chrome I think, could also have been safari. Using the original TeX source for the screen reader is significantly hindered by the fact that TeX isn't a markup-, but a programming language. TeX is turing-complete by design, making it infinitely extensible, in order to avoid knuth ever having to re-typeset his books. After all if it's a programming language and a new system, problem, style, e.t.c, comes along you can just write a program that deals with it. But this has horrible effects on render-ability, in order to know what the final document should look like, you need to run it, no way around that, thanks to Rice's theorem. LaTeX users also generate a lot of their figures with TeX itself, write their own styling or bibliography rules, and write their own custom graphics rendering libraries. The easiest path is to just give up, render the entire PDF to a 300dpi lossless image, and throw it into an OCR engine. These things contain a buttload of heuristics to generate structure and meaningful text based on visuals, and since we know that humans explicitly (and often only) care about those in TeX documents... It's pretty darn sad, because in many ways TeX is holding scientific advances back, by eating its own children. My guess is that if scientific papers had branched off of plaintext tools like troff, we'd probably publish papers as machine readable semantically annotated knowledge-graphs by now.
- marvy 6y agoI just tried in Chrome; at least for page 2 it actually did a bit better than my browser: it actually managed to preserve the line breaks! (I don't have Safari installed.) (Also, if you look closely, the summation signs are not gone, they are replaced by the letter P. Which is not helpful I admit.) Do you think there's any place here for education/advocacy? For instance, everyone who makes web pages knows to provide alt text for images. If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right? I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR. (Though OCR is still an improvement over the best option in the 1980s I bet. Though I wasn't around so just guessing.)
- j-pb 6y ago"Do you think there's any place here for education/advocacy? For instance, everyone who makes web pages knows to provide alt text for images." That's not a culture thing, this is by the mechanism of <img alt. You won't get people who write TeX to change all of their workflows, and even if there was such a culture. "If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right?" Then it would still be nigh impossible because TeX commands, like all programming languages, compose rather poorly. It would be a herculean effort to produce a kinda but not really TeX that is both accessible with a focus on semantics, yet still compatible with the billions of lines of LaTeX/TeX out there. "I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR." Boycott LaTeX/TeX and PDF everywhere you can. Whenever you publish a paper, also publish it in markdown/html. Publish in OpenAccess Journals like [PeerJ](https://peerj.com/ https://peerj.com/) which convert all of their papers to html in addition to pdf. Consider publishing papers in alternative forms like nextjournal.com . We need to get our priorities straight in academia :/. This entire "but latex produces such beautiful documents", "I'm working towards getting into the most prestigious journal" culture of snobbery and vanity needs to stop. We need to go back to caring about the content, not the presentation, something TeX ironically was meant to do.