7 ms·
HTML as an Accessible Format for Papers (2023)
- el3ctron 10mo agoAccessibility barriers in research are not new, but they are urgent. The message we have heard from our community is that arXiv can have the most impact in the shortest time by offering HTML papers alongside the existing PDF.
- lalithaar 10mo agoHello, I was going through html versions of my preprints on Arxiv, thank you for all that you guys do Please do let me know if the community could contribute through any means for the same
- dginev 10mo agoYou can help make LaTeXML better, or you can simply report issues when you spot them during reading. Some we have collected automatically (any errors and missing packages), but others we can't - wrong colors, broken aspect ratios of figures, weirdly layed out author lists, etc.
- lalithaar 10mo agoI was reading through this article too, glad to have found it on here
- ForceBru 10mo agoIs this new or somehow updated? HTML versions of papers have been available for several years now. EDIT: indeed, it was introduced in 2023: https://blog.arxiv.org/2023/12/21/accessibility-update-arxiv-now-offers-papers-in-html-format/ https://blog.arxiv.org/2023/12/21/accessibility-update-arxiv...
- Tagbert 10mo agoFrom the paper... Why "experimental" HTML? Did you know that 90% of submissions to arXiv are in TeX format, mostly LaTeX? That poses a unique accessibility challenge: to accurately convert from TeX—a very extensible language used in myriad unique ways by authors—to HTML, a language that is much more accessible to screen readers and text-to-speech software, screen magnifiers, and mobile devices. In addition to the technical challenges, the conversion must be both rapid and automated in order to maintain arXiv’s core service of free and fast dissemination.
- ForceBru 10mo agoNo I mean _arXiv_ has had experimental support for generating HTML versions of papers for years now. If you visit arXiv, you'll see a lot of papers have generated HTML alongside the usual PDF, so I'm trying to understand whether the article discussed any new developments. It seems like it's not new at all
- daemonologist 10mo agoThere are pretty often problems with figure size and with sections being too narrow or wide (for comfortable reading). The PDF versions are more consistently well-laid-out.
- fooofw 10mo agoIt's kind of fun to compare this formulation with the seemingly contradictory official arXiv argument for submitting the TeX source [1]: > 1. TeX has many advantages that make it ideal as a format for the archives: It is plain text, it is compact, it is freely available for all platforms, it produces extremely high-quality output, and it retains contextual information. > 2. It is thus more likely to be a good source from which to generate newer formats, e.g., HTML, MathML, various ePub formats, etc. [...] Not that I disagree with the effort and it surely is a unique challenge to, at scale, convert the Turing complete macro language TeX to something other than PDF. And, at the same time, the task would be monumentally more difficult if only the generated PDFs were available. So both are right at the same time. [1] https://info.arxiv.org/help/faq/whytex.html#contextual https://info.arxiv.org/help/faq/whytex.html#contextual
- inglor 10mo agoYou're right https://github.com/arXiv/arxiv-docs/blob/develop/source/about/accessible_HTML.md https://github.com/arXiv/arxiv-docs/blob/develop/source/abou... this needs a 2023 tag @dang
- ashleyn 10mo agoCan't help but wonder if this was motivated in part by people feeding papers into LLMs for summary, search, or review. PDF is awful for LLMs. You're effectively pigeonholed into using (PAYING for) Adobe's proprietary app and models which barely hold a candle to Gemini or Claude. There are PDF-to-text converters, but they often munge up the formatting.
- jrk 10mo agoNot sure when you last tried, but Gemini, Claude, and ChatGPT have all supported pretty effective PDF input for quite a while.
- Barbing 10mo ago>Did you know that 90% of submissions to arXiv are in TeX format, mostly LaTeX? That poses a unique accessibility challenge: to accurately convert from TeX—a very extensible language used in myriad unique ways by authors—to HTML, a language that is much more accessible to screen readers and text-to-speech software, screen magnifiers, and mobile devices. Challenging. Good work!
- sega_sai 10mo agoUnfortunately I didn't see the recommendation there on what can be done for old papers. I checked, and only my papers after 2022 have an HTML version. I wish they'd make some kind of 'try html' button for those.
- sundarurfriend 10mo agoDo the older papers work via [Ar5iv](https://ar5iv.labs.arxiv.org/ https://ar5iv.labs.arxiv.org/) ? > View any arXiv article URL [in HTML] by changing the X to a 5 The line > Sources upto the end of November 2025. sounds to me like this is indeed intended for older articles.
- dginev 10mo agoar5iv tracks the arXiv collection with a one month lag. Exactly as to signal that this is not the "official" arXiv rendering. It is also a showcase predating the arXiv /html/ route, but largely using the same technology. Nowadays maintained by the same people (hi!) There used to be another showcase, called arxiv-vanity. They captured what happened pretty well with their farewell post on their homepage: https://www.arxiv-vanity.com/ https://www.arxiv-vanity.com/
- rootnod3 10mo agoMaybe unpopular, but papers should be in n markdown flavor to be determined. Just to have them more machine readable.
- doc_ick 10mo agoNot unpopular, but a lot of the publishing companies would have to agree to that and make their own formatting/structure rules. I also haven’t had good luck with images/graphs/custom tables in anything but typist/latex.
- xigoi 10mo agoCompared to HTML, Markdown is very bad at being mahcine-readable.
- nateroling 10mo agoSeeing the Gemini 3 capabilities, I can imagine a near future where file formats are effectively irrelevant.
- doc_ick 10mo agoTell that to publishing companies.
- DANmode 10mo agoFiles. Truth in general, if we aren't careful.
- sansseriff 10mo agoSeriously. More people need to wake up to this. Older generations can keep arguing over display formats if they want. Meanwhile younger undergrad and grad students are getting more and more accustomed to LLMs forming the front end for any knowledge they consume. Why would research papers be any different.
- JadeNB 10mo ago> Meanwhile younger undergrad and grad students are getting more and more accustomed to LLMs forming the front end for any knowledge they consume. Well, that's terrifying. I mean, I knew it about undergrads, but I sure hoped people going into grad school would be aware of the dangers of making your main contact with research, where subtle details are important, through a known-distorting filter. (I mean, I'd still be kinda terrified if you said that grad students first encounter papers through LLMs. But if it is the front end for all knowledge they consume? Absolutely dystopian.)
- sansseriff 10mo agoI admit it has dystopian elements. It’s worth deciding what specifically is scary though. The potential fallibility or mistakes of the models? Check back in a few months. The fact they’re run by giant corps which will steal and train on your data? Then run local models. Their potential to incorporate bias or persuade via misalignment with the reader’s goals? Trickier to resolve, but various labs and nonprofits are working on it. In some ways I’m scared too. But that’s the way things are going because younger people far prefer the interface of chat and question answering to flipping through a textbook. Even if AI makes more mistakes or is more misaligned with the reader’s intentions than a random human reviewer (which is debatable in certain fields since the latest models game out), the behavior of young people requires us to improve the reputability of these systems. (Make sure they use citations, make sure they don’t hallucinate, etc). I think the technology is so much more user friendly that fixing the engineering bugs will be easier than forcing new generations to use the older systems.
- jas39 10mo agoPandoc can convert to svg. It can then be inlined in html. Looks just like latex, though copy/paste isn't very useful
- stephenlf 10mo agoThat doesn’t solve the accessibility issue, though. You need semantic tags.
- sundarurfriend 10mo ago[Sept 2023] as per the wayback machine.
- billconan 10mo agoI don't think HTML is the right approach. HTML is better than PDF, but it is still a format for displaying/rendering. the actual paper content format should be separated from its rendering. i.e. it should contain abstract, sections, equations, figures, citations etc. but it shouldn't have font sizes, layout etc. the viewer platforms then should be able to style the content differently.
- afavour 10mo agoWouldn’t that be CSS?
- billconan 10mo agono <div class="abstract-container"> <div class="abstract"> <pre><code> abstract text ... </code></pre> </div> <div class="author-list"> <ol> <li>author one</li> <li>author two</li> <ol> </div> should be just: [abstract] abstract text [authors] author one | email | affiliation author two | email | affiliation
- afavour 10mo agoSounds like XML and XSL would be a great fit here. Shame it’s being deprecated. But you could still use HTML. Elements with a dash in are reserved for custom elements (that is, a new standardised element will never take that name) so you could do: <paper-author-list> <paper-author /> </paper-author-list> And it would be valid HTML. Then you’d style it with CSS, with paper-author { display: list-item; } And so on.
- bawolff 10mo agoNothing is stopping you from using server side XSL. I personally dont think its a great fit, but people need to stop acting like xsl has been wiped from the face of the earth.
- 10mo ago
- vatsachak 10mo agoWhy do we like HTML more than pdfs? HTML rendering requires you to be connected to the internet, or setting up the images and mathJax locally. A PDF just works. HTML obviously supports dynamic embedding, such as programs, much better but people just usually post a github.io page with the paper.
- devnull3 10mo ago> HTML rendering requires you to be connected to the internet Not really. One can always generate a self-contained html. Both CSS and JS (if needed) can be inline.
- vatsachak 10mo agoTrue but the webdev idiom is injecting things such as mathjax from a cdn. I guess one can pre-render the page and save that, but that's kind of like a PDF already
- recursive 10mo agoWhy would html rendering require a network connection? It doesn't seem to on my machine.
- teddy-smith 10mo agoIt's extremely easy to convert HTML/CSS to a PDF with the print to PDF feature of the browser. All papers should be in HTML/CSS or Tex then just simply converted to PDF. Why are we even talking about this?
- tefkah 10mo agoWhat are you talking about? No one’s writing their paper in HTML. The problem is having the submissions be in TeX and converting that to HTML, when the only output has been PDF for so long. The problem isn’t converting HTML to PDF, it’s making available a giant portion of TeX/pdf only papers in HTML. If you’re arguing that maybe TeX then shouldn’t be the source format for papers then I agree, but other than Typst (which also isn’t perfect about HTML output yet) there aren’t that many widely accepted/used authoring formats for physics/math papers, which is what ArXiV primarily hosts.
- teddy-smith 10mo agoThis is what I'm talking about. HTML/CSS is more powerful than PDF or TEX. https://csszengarden.com/ https://csszengarden.com/
- nkrisc 10mo agoSo, uh, where do the HTML versions of the papers come from?
- teddy-smith 10mo agoGround truth.
- nkrisc 10mo agoWhat do you mean by that? That researchers should be authoring their papers in HTML?
- ekjhgkejhgk 10mo agoLOL what. You're either trolling, or you've never written a paper in your life.
- cubefox 10mo agoThis is not new, the title should say (2023). They have shipped the HTML feature with "experimental" flag for two years now, but I don't know whether there is even any plan to move out of the experimental phase. It's not much of an "experiment" if you don't plan to use some experimental data to improve things somehow.
- ekjhgkejhgk 10mo agoI wish epub was more common for papers. I have no idea if there's any real difficulties with that, or just not enough demand.
- pspeter3 10mo agoWhy epub? Isn’t it just HTML under the hood?
- ekjhgkejhgk 10mo agoBecause I can open it on my ereader.
- silon42 10mo agoI think it should also have JS disabled (I hope!)
- mmooss 10mo agoepub is html, under the hood Is there an epub reader that can format text approximately as usably and beautifully as pdf? What I've seen makes it noticeably harder to read longer texts, though I haven't looked around much. epub also lacks annotation, or at least annotation that will be readable across platforms and time.
- hombre_fatal 10mo agoBecause what makes epub a format on top of html is just that someone QA'ed it and wrote the html/css with it in mind. Especially considering things like diagrams and tables. Not really what you want researchers to waste their time doing. But you can use any of the numerous html->epub packagers yourself.
- leobg 10mo agoIt must have been around 1998. I was editor of our school’s newspaper. We were using Corel Draw. At some point, I proposed that we start using HTML instead. In the end, we decided against it, and the reasons were the same that you can read here in the comments now.
- DominikPeters 10mo agoAs an arXiv author who likes using complicated TeX constructions, the introduction of HTML conversion has increased my workload a lot trying to write fallback macros that render okay after conversion. The conversion is super slow and there is no way to faithfully simulate it locally. Still I think it's a great thing to do.
- xworld21 10mo agoI believe dginev's Docker image https://github.com/dginev/ar5ivist https://github.com/dginev/ar5ivist is very close to what runs on arXiv and can be run locally. It uses a recent LaTeXML snapshot from September.
- _dain_ 10mo agoWasn't the World Wide Web invented at CERN specifically for sharing scientific papers? Why are we still using PDFs at all?
- fsh 10mo agoNo, it wasn't. Scientists at CERN used DVI and later PDF like everyone else. HTML has no provisions for typesetting equations and is therefore not suitable for physics papers (without much newer hacks such as MathML).
- teddy-smith 10mo agoWhy not typeset in something else and import the image into html/css?
- cxr 10mo agoMathML isn't new. It predates Windows 98 and the birth of a substantial part of HN's userbase.
- ComputerGuru 10mo agoIf the Unicode consortium would spend less time and effort on emoji and more on making the most common/important mathematical symbols and notations available/renderable in plain text, maybe we could move past the (LA)TeX/PDF marriage. OpenType and TrueType now (edit: for well over a decade, actually) support the necessary conditional rendering required to perform complicated rendering operations to get sequences of Unicode code points to display in the way needed (theoretically, anyway) and with fallback missing-glyph-only font family substitution support available pretty much everywhere allowing you to seamlessly display symbols not in your primary font from a fallback asset (something like Noto, with every Unicode symbol supported by design, or math-specific fonts like Cambria Math or TeX Gyre, etc), there are no technical restrictions. I’ve actually dug into this in the past and it was never lack of technical ability that prevented them from even adding just proper superscript/subscript support before, but rather their opinion that this didn’t belong in the symbolic layer. But since emoji abuse/rely on ZWJ and modifiers left and right to display in one of a myriad of variations, there’s really no good reason not to allow the same, because 2 and the squares symbol are not semantically the same (so it’s not a design choice). An interesting (complete) tangent is that Gemini 3 Pro is the only model I’ve tested (I do a lot of math-related stuff with LLMs) that absolutely will not under any circumstances respect (system/user) prompt requests to avoid inline math mode (aka LATeX) in the output, regardless of whether I asked for a blanket ban on TeX/MathJax/etc or when I insisted that it use extended unicode codes points to substitute all math formula rendering (I primarily use LLMs via the TUI where I don’t have MathJax support, and as familiar as I once was with raw TeX mathematical notations and symbols, it’s still quite easy to confuse unrendered raw output by missing something if you’re not careful). I shared my experiment and results here – Gemini 3 Pro would insist on even rendering single letter constants or variables as $k$ instead of just k (or k in markdown italics, etc) no matter how hard I asked it not to (which makes me think it may have been overfit against raw LATeX papers, and is also an interesting argument in favor of the “VL LLMs are the more natural construct”): https://x.com/NeoSmart/status/1995582721327071367?s=20 https://x.com/NeoSmart/status/1995582721327071367?s=20
- raincole 10mo agoMath formulas are far far far more complex than unicode emojis. I don't even know how to start comparing them.
- percentcer 10mo agoDumb question but what stops browsers from rendering TeX directly (aside from the work to implement it)? I assume it's more than just the rendering
- pwdisswordfishy 10mo agoFor starters, TeX is Turing-complete, and the tokenizer is arbitrarily reprogrammable at runtime.
- ErroneousBosh 10mo agoOkay then, what would stop you rendering TeX to SVG and embedding that? Edit: Genuine question, not rhetorical - I don't know how well it would work but it sounds like it should.
- fooofw 10mo agoThat would (mostly if not always) work in the sense of reproducing the layout of the pages, but would defeat the purpose of preserving the semantic information present in the TeX file (what is a heading, a reference and to what, a specific math environment, etc.) which is AFAIK already mostly dropped on conversion to PDF by the latex compiler.
- ErroneousBosh 10mo agoCouldn't you write a TeX renderer that emitted HTML (or RST, or Markdown, or whatever) with SVG for the equations?
- fooofw 10mo agoI think this project is based on LaTeXML (https://math.nist.gov/~BMiller/LaTeXML/ https://math.nist.gov/~BMiller/LaTeXML/) which is exactly that (except for the SVG part)
- gbear605 10mo ago
- dginev 10mo agoHi, an arXiv HTML Papers developer here. As a very brief update - we are pending a larger update. You will spot many (many) issues with our current coverage and fidelity of the paper rendering. When they jump at you, please report them to us. All reports from the last 2 years have landed on github. We have made a bit of progress since, but there are (a lot of) more low-hanging fruit to pick. Project issues: https://github.com/arXiv/html_feedback/issues/ https://github.com/arXiv/html_feedback/issues/ The main bottleneck at the moment is developer time. And the main vehicle for improvements on the LaTeX side of things continues to be LaTeXML. Happy to field any questions.
- istillwritecode 10mo agoI would like to write code for latexml to translate a package but I found the documentation to be hard to understand. That might be what is holding developers back. I looked at this a year ago and gave up.
- dginev 10mo agoTell us what you would need described in a tutorial to be productive, as well as your background with the technologies involved (TeX/LaTeX, perl, XML, XSLT, HTML). Probably best as a new issue: https://github.com/brucemiller/LaTeXML/issues https://github.com/brucemiller/LaTeXML/issues It's a pretty deep rabbit hole, but I wholeheartedly agree most standard package support incantations should be easy and few to use.
- chr15m 10mo agoWish I could upvote this harder. Thank you arXiv!
- RandyOrion 10mo agoFor arXiv papers, I prefer HTML format much more than PDF format. Compared to PDF format, HTML format is much more accessible because of browsers. Basically I can reuse my browser extensions to do anything I like without hassle, like translation, note taking, sending texts to LLMs, and so on. For now, arXiv offers two HTML services: the default one in https://arxiv.org/html/xxxx.xxxxx https://arxiv.org/html/xxxx.xxxxx , and the alternative one in https://ar5iv.labs.arxiv.org/html/xxxx.xxxxx https://ar5iv.labs.arxiv.org/html/xxxx.xxxxx , here 'x' is a placeholder for a number or digit. The most glaring problem of the default HTML service is the coverage of papers. Sometimes it just doesn't work, e.g., https://arxiv.org/html/2505.06708 https://arxiv.org/html/2505.06708 . The solution may be switch to alternative HTML service, e.g., https://ar5iv.labs.arxiv.org/html/2505.06708 https://ar5iv.labs.arxiv.org/html/2505.06708 . Note that alternative HTML service also has coverage problem. Sometimes both HTML services fail, e.g. https://arxiv.org/abs/2511.22625 https://arxiv.org/abs/2511.22625 .
- rhubarbtree 10mo agoSerious question: do websites from the 90s work well in modern browsers? Because PDFs from that time view fine.
- cxr 10mo agoAside from sites that used non-standard stuff like ActiveX or Java applets, the general answer is "yes". And to respond to your implied criticism: the stability/reliability/fidelity of PDFs is a myth. It would be hard to say how many dozens of PDFs I've come across in the last two years that don't look the same across devices/viewers (or sometimes just fail to render in their entirety). This played a significant part in a cascade of errors in one incident I know of that resulted in the payout of a claim more than $1,000 but less than $10,000—not to mention a lot of strife and anger for the persons involved over the course of multiple months before resolution. (As I write this now, I realize I'd almost forgotten about the fact that almost every time I've taken something to FedEx or UPS to be printed at a self-service kiosk, the result has been unusable, so I've had to take it to the clerk to have them print it instead.) HTML at least has the property that it's still trivial to access and extract the data if you run into either malformed inputs or ones that are valid but incompatible/unsupported by whatever viewer (browser) you happen to be using, which is a lot more than you can say for more opaque formats like Java, PDF, and Flash.
- notorandit 10mo agoThee problem is the viewer, not the format. We are talking about accessibility and scientific papers, where fancy animations and transitions are not core features. LaTeX and TeX are the de facto standard for this context and converting all existing documents is a lot of work and energy to be spent for basically little gain, if any.
- constantcrying 10mo agoReading this thread many people do not seem to understand what to the problem even is. What researchers writing Papers want is a low effort/high flexibility way to write documents (Nobody wants to write their paper in HTML). For a paper to be printed it needs to be in some printable format, like PDF. To provide accessibility and accommodate the changing ways papers are read, which is increasingly online, HTML is also a desirable output. What really is needed is a markup language which natively can target both PDF and HTML. This is something typst is working on, but I am not aware of any other project, which either comes close to the features of LaTeX or supports both target formats. To me this is the only reasonably way to address the accessibility and usability issues around Papers. Have one markup, with sufficient accessibility features, which simultaneously targets HTML and PDF.
- zipy124 10mo agoThe biggest issue with papers for me today is that they don't allow videos as anything other than supplemental materials to be downloaded, or linking to a web-page that has them. I want to embed gif's or videos in my papers directly!
- cxr 10mo agoHere in the Muggle world, there's no material I know of that can be used to produce a type of paper that supports moving images.