3 ms·
Over the years many people have hypothesized that once WASM was really mature, it would become practical to fix the issues with web browser layout by sending do
by jerf 3mo ago
Over the years many people have hypothesized that once WASM was really mature, it would become practical to fix the issues with web browser layout by sending down custom layout machines to users.
I would find it hilarious if LaTeX turned into a leader in that space. I doubt it could hold on to that. There's a lot of things that something designed from the beginning for web-like uses could probably improve on that would be capable of overcoming LaTeX. But I could see a world where it carves out a niche and holds on to that niche for a long period of time.
- nicoburns 3mo agoRunning layout in WASM is already practical. A good demo is https://www.nicbarker.com/clay https://www.nicbarker.com/clay The things you can't do are things like expose an accessibility tree (without a dummy DOM), interact with the system IME, and access system fonts.
- jerf 3mo agoI feel like it's fair to say that you have not "fixed the issues with browser layout" if you lose accessibility and input. System fonts I can live without, we can push our own, but those two things are a big deal. Even input you might be able to hack around but accessibility is a big deal and the "hack" at that point is nearly to both lay it out in the browser and the supposed "fixed" layout system, and while that may work in some sense I again have lots of questions about whether that is really "fixed".
- nicoburns 3mo agoI mean, I agree it's not fixed. I guess I'm just saying it's not the layout engine that's the blocker. FWIW, I actually think it would be much more valuable just to fix the spec and make CSS layout fast-by-default.
- exe34 3mo agoNowadays I imagine OCR and vllms would solve that? Tesseract is incredibly fast and accurate.
- jerf 3mo agoAccessibility is not just about getting text out. It's about the navigation flow, integration into various OS features, Aria attributes in HTML and what they do. Plus you may not count this as "accessibility" but things like integrating fully with multi-lingual text entry methods. It's solvable, except for what isn't solvable because the browser doesn't expose it, which can be solved by fixing that too. But it's a lot more work than meets the eye. And building the layout engine is hard enough in the first place. Give me a week and an AI budget and I'll produce you some sort of layout engine that works when fed exactly the sorts of inputs I anticipated, but to build something that survives contact with the real world is going to be well beyond something you just prompt your way around today.
- gcr 3mo agoThis page starts flickering madly when I pinch-to-zoom. Until a11y details like this are figured out, I don’t think this should be considered for general use beyond a cool prototype.
- PaulDavisThe1st 3mo agoI added DVI support to NCSA Mosaic back in 1993-94, believing it to be a better format for "rich" documents than HTML or PDF. Nobody else seemed convinced :(
- gjvc 3mo agopretty sure I remember reading about this with excitement and wonder ...
- dhosek 3mo agoThe problem with DVI is twofold: First, font support is purely by reference which means that you need some way of connecting the fonts used in the document with the DVI file. Use of the wrong font could produce some spectacularly bad output. Second, graphical support, other than rectangular boxes is only handled through the xxx opcode which never had any standardized meaning (although I tried). This limitation also applied to colors. Really, it was only with the final victory of PDF as the universal document format that these limitations were finally ameliorated.
- PaulDavisThe1st 3mo agoI would never advocate for DVI today. In 1993/94 ? It was a much stronger candidate, and given that any DVI (including images) could be rendered into PDF (or whatever), its supposed technical limitations at that time were not of much significance.
- stereo 3mo agoHtml back then had the same limitations.
- dhosek 3mo agoNot really. <img> was a first-class tag if not from the beginning, pretty early on and the same with color.
- LudwigNagasena 3mo agoHard to imagine anything worse than LaTeX for web layout. Imagine resizing a page and waiting for the re-compilation of the whole page.
- jerf 3mo agoThat's part of the reason I'd find it so funny, yes. The reason why I consider it a possibility is that LaTeX has two things out of the gate: The technical capability, and a small but arguably rabid user base. It's the sort of thing that can take an early lead but is quite unlikely to sustain it. But you can't deny that LaTeX has had incredible staying power, despite the list of issues that everyone who uses it has with it.
- Onavo 3mo agoThere's also PDF/PostScript as a format, but it's also made for fixed/absolute/print layouts. HTML/CSS (and I suppose related tech like XAML) were the only technologies where responsiveness was first priority. Now if only the browser conglomerates can disrupt Pantone too...
- jampekka 3mo agoI use LaTeX daily and hate it with a passion. I kinda lost all hope when ArXiv decided to do the HTML support by hacking a LaTeX to HTML conversion. We already have a very powerful layouting engine: the web browser. The only missing piece is printing/pagination, for which there was some CSS Paged Media progress, but that stalled. However, why the hell are we even doing paged media? For screen viewing it's strictly worse, and very few people print papers anymore. And even for those HTML pages print passably enough.
- rwl 3mo agoCitations. Until we can cite specific passages in HTML as quickly, easily and readably as we can cite pages in paged media, HTML will remain a second class citizen for serious scholarship.
- jampekka 3mo agoSections, anchors, url text fragments. In many fields citing a page is also quite rare.
- rwl 3mo agoYes, I'm aware. The problem with anchors is that the author has to take care to make them both unique and human-readable across the whole document. The problem with URL text fragments is that they generally become too long to be a citation. Both need quality-of-life improvements in the tooling before academics will consider using them. For example, it's pretty easy to use CSS counters to create as many automatically incrementing counters as you need. Page number, section number, theorem number, etc. But (a) you can't do this in raw HTML, which means you can't rely on this working in any user agent that has poor support for CSS; and (b) you can't refer to the values of those counters outside of CSS. Thus, there's no way for you, as an author, to write the write the equivalent of LaTeX's "See Theorem~\ref{thm:foo}" in HTML and rely on having it auto-numbered and rendered correctly wherever your readers are. So the number has to already be there as text in the HTML, and you need a separate, document-unique id attribute coordinated with the number to use for linking. This means you need a separate compilation step to produce the HTML, so you've lost pretty much the only advantage that HTML had over LaTeX. But fundamentally, the problem is that citation practices assume that the reading format is controlled by the author/publisher, so that all readers are looking at a common view of the document, whereas the Web assumes that the reading format is controlled by the user agent, and user agents vary widely. Thus, on the Web, you need formatting-independent citation practices, which have not yet evolved, or at least not become widely used, because it's too hard compared to page-based media, which have been working for centuries.