4 ms·
I just tried in Chrome; at least for page 2 it actually did a bit better than my browser: it actually managed to preserve the line breaks! (I don't have Safari
by marvy 6y ago
I just tried in Chrome; at least for page 2 it actually did a bit better than my browser: it actually managed to preserve the line breaks! (I don't have Safari installed.)
(Also, if you look closely, the summation signs are not gone, they are replaced by the letter P. Which is not helpful I admit.)
Do you think there's any place here for education/advocacy? For instance, everyone who makes web pages knows to provide alt text for images.
If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right?
I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR.
(Though OCR is still an improvement over the best option in the 1980s I bet. Though I wasn't around so just guessing.)
- j-pb 6y ago"Do you think there's any place here for education/advocacy? For instance, everyone who makes web pages knows to provide alt text for images." That's not a culture thing, this is by the mechanism of <img alt. You won't get people who write TeX to change all of their workflows, and even if there was such a culture. "If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right?" Then it would still be nigh impossible because TeX commands, like all programming languages, compose rather poorly. It would be a herculean effort to produce a kinda but not really TeX that is both accessible with a focus on semantics, yet still compatible with the billions of lines of LaTeX/TeX out there. "I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR." Boycott LaTeX/TeX and PDF everywhere you can. Whenever you publish a paper, also publish it in markdown/html. Publish in OpenAccess Journals like [PeerJ](https://peerj.com/ https://peerj.com/) which convert all of their papers to html in addition to pdf. Consider publishing papers in alternative forms like nextjournal.com . We need to get our priorities straight in academia :/. This entire "but latex produces such beautiful documents", "I'm working towards getting into the most prestigious journal" culture of snobbery and vanity needs to stop. We need to go back to caring about the content, not the presentation, something TeX ironically was meant to do.
- marvy 6y agoSo basically, give up on LaTeX to PDF as the primary workflow. Either convert LaTeX to semantic HTML instead of to PDF (probably doable for simple cases... \emph to <em>, \section to <h2>, \begin{tabular) to <table>, and so on), or better yet just author HTML without going through LaTeX at all. If needed, convert the semantic HTML to PDF as well, for printing or whatever, and then the PDF might end up sane. Fine, fair enough. If fixing TeX is hopeless, then so be it. I assumed it just needed a few small tweaks, maybe combined with slightly cleverer PDF viewers. Guess I was wrong. But then what should people use for math? I suppose there's MathJax, which seems to have put some thought into accessibility. There's still a problem though. I can't help but notice that the journal you linked to is a biology journal. In some math/CS circles which are TeX's "home turf", TeX is far more entrenched to the point where I'm not sure such things even exist. For instance, arxiv sort of supports HTML, but not really: https://arxiv.org/help/submit_html https://arxiv.org/help/submit_html So there I guess step 1 is to make HTML a viable option.
- j-pb 6y agoPeerJ has actually quite a few publications, one of them is CS ;) https://peerj.com/computer-science/ https://peerj.com/computer-science/ the dropdown on the top left allows you to switch between them. I think MathJax is certainly a step in the right direction, they even support rendering to MathML. But I agree that there is a certain lack there in terms of full semantic representations. MathJax is more accessible than TeX but it's still describing visual layout, instead of semantic meaning. Pushing HTML to arxiv is also a step into the right direction. I think the most important thing we can do is not be complacent with the state of the art. We need to go back to an age of computing where we didn't think we had it all figured out. We need to experiment, and not be afraid to take a step back in some aspects, like layout and kerning, in exchange for other advances like semantic representations and knowledge representation. I think bred victor has a great talk on this: https://www.youtube.com/watch?v=8pTEmbeENF4 https://www.youtube.com/watch?v=8pTEmbeENF4 I think we need to experiment with things like observablehq.com or nextjournal.com or the many other that are coming into existence.
- 6y ago