11 ms·
Converting my PhD thesis into HTML (2021)
- bravura 4y agoI think the simplest solution is uploading your thesis to arxiv.org, then using arxiv-vanity (based upon LaTeXML) to render your arxiv link as a responsive web page.
- jimhefferon 4y agoLaTeXML does great things, but it also has limitations. A doc that is in generic LaTeX is going to process much better than one with significant customizations. But it is a good tool for sure.
- 131hn 4y agoThe thesis french transcript of the tl;dr is written in verse/alexendrin. Because, why not
- 082349872349872 4y agoexcellent !
- breck 4y agoIf I was writing a PhD thesis today, I'd use Scroll: https://scroll.pub/ https://scroll.pub/
- amelius 4y agoThe FAQ does not explain what scroll is or does near the top of the document. It seems to produce HTML only.
- breck 4y agoBetter printable PDF support is coming.
- CJefferson 4y agoPlease tell us when you are self-promoting your own stuff. I would currently recommend against any PhD student using it for their PhD -- it can't even generate a PDF with required formatting, which is required by most Universities.
- breck 4y ago> Please tell us when you are self-promoting your own stuff. Sorry! I should have said our. You are right to call me out on that. Sometimes I don't reread my comments before posting. > I would currently recommend against any PhD student using it for their PhD -- it can't even generate a PDF with required formatting, which is required by most Universities. If it doesn't do something someone needs, file an issue or email me and we can probably add what is needed.
- jltsiren 4y agoPDF documents have two main benefits. The entire document is a single file, and we know that old documents work. I regularly read papers from 10 or 20 years ago and sometimes even ones from 30 years ago. The old PDF documents work without any major issues. I have much less confidence that future browsers will continue displaying old HTML documents laid out using then-obsolete techniques in a sane way on future hardware.
- felixfbecker 4y agoWhat HTML document from 10 or 20 years ago does not still render in a modern browser? Modern browsers are extremely good to maintaining backwards compatibility at all costs (aka "don't break the web"), to the degree many on HN often argue it hinders evolving web technologies.
- jltsiren 4y agoThe ones that relied on Flash, for example. And the layouts designed for 800x600 displays may not work particularly well on modern computers.
- pclmulqdq 4y agoI originally wanted to blog using LaTeX, and convert that to HTML. I ran into all of these options, and also started writing my own LaTeX->HTML flow, but it got too complicated to be a hobby project. It turns out that there are parts of LaTeX that are really hard to generically convert to web constructs. I settled for markdown with KaTeX for math, although I would like to return to the LaTeX->HTML project at some point soon.
- aitchnyu 4y agoIn The Art of Unix Programming (2003) the authors assert simple text formats can be grepped, awked and are easy to compose in a text editor. Hand typing xml is cruel. But now text editors have perfect syntax highligting and squiggles to pinpoint errors, autocomplete, automatic formatting and toolchains to eliminate errors and extract info. Isnt html or a high level abtraction (like Spectacle) the best tool for the job today? https://formidable.com/open-source/spectacle/docs/#one-html-page https://formidable.com/open-source/spectacle/docs/#one-html-...
- pjmlp 4y agoHence why one should use stuff like https://www.oxygenxml.com/ https://www.oxygenxml.com/ Similar products have been in business since the early 2000's.
- vouaobrasil 4y agoI disagree with the author that PDFs are a terrible format. They guarantee layout, which is very important for complex scientific presentations. Even slight differences in layout can make a complex set of equations difficult to parse. LaTeX also has a much superior word-break/hyphening algorithm to the HTML engines of browsers. I find PDF math papers easy to browse, unlike the author. They're much easier and more organized than a website, can be easily searched and have a *proper table of contents* compared to websites. As for poorly browsable on a phone -- well I think that is irrelevant because nobody is going to read a complex technical paper in practise on a phone. They do look decent in tablets, and as for screen readers...well that's a valid point but screen readers don't work well for material with lots of equations anyway. I applaud the author for the effort but looking at the result, I would not want to read math that way.
- jech 4y ago> I find PDF math papers easy to browse So do I. Still, I wish LaTeX produced easily reflowable PDFs, especially when a document is formatted in two columns.
- enriquto 4y agoBut it does, doesn't it? You add the "twocolumn" option and recompile. Unless your LaTeX is too fancy this will tipically give a very good result (at worst, some figures with hardcoded sizing will be awkardly placed).
- jech 4y agoI cannot do that when I'm reading a paper written by somebody else, and I only have the produced PDF.
- abdullahkhalids 4y agoThat's why arxiv is a god send, because the source is available there, if the author has uploaded it there. Science needs a culture of open sharing, the same way physics and math has it.
- dwheeler 4y agoThe problem here is specific to LaTeX. I wrote my PhD dissertation using OpenOffice.org (now use LibreOffice), and generating HTML was easy (I posted the HTML). But the author is right, LaTeX is widely used, translating it to HTML is hard, and there are no incentives to make or improve tools. Even if you don't want HTML, it'd be good for the LaTeX tools to automaticalky generate reflowable PDF for accessibility. There should be a process for funding infrastructure to accelerate science, and this would be a good example. There's an interesting trick you could try. PDF supports embedding other contents. LibreOffice, for example, can slip its original edited file into a generated PDF, producing perfectly editable PDF. Maybe a variant of this idea could be used, e.g., store the LaTeX source of HTML in the PDF, so people can "get the PDF" yet still have options. But that's just a side idea, the real issue is funding infrastructure for science.
- V1ndaar 4y agoCurrently finishing up my own PhD thesis. My approach to the same problem is quite different. I write my thesis in Org mode. Exporting to HTML is pretty painless. Been doing the same for years for my notes. PDF export via LaTeX & HTML export. LaTeX and PDFs fail pretty hard when including source code (some literate programming in Org). That was my initial motivation behind also producing HTML. The final thesis that I will hand in is of course a regular PDF (well, a print based on that). But the HTML version can contain lots more stuff that doesn't fit (and belong) into the actual paper thesis, e.g. code snippets to generate plots etc. (optional export of Org subsections). By publishing the git repository of the thesis, linking all code and data + a bit of work -> full reproducible thesis.
- hoosieree 4y agoHeh, I also write papers in org and am currently writing my dissertation in org. Source code is always a pain to export for PDF, especially when switching from 1 to 2 column layout depending on the publication. My blog is written in org too, but I post-process to make it fit in with the rest of my static site. At some point maybe I'll get enough free time to swap out my makefile setup for org-publish, but if it ain't broke... To anyone who'll listen I advocate for org-mode as a better alternative to Jupyter notebooks, Markdown, and LaTeX. It's in some ways the antithesis to "do one thing well". If you try to do N things well while adhering to the unix philosophy you end up learning N different tools. But org-mode is one tool that does N things well, and some of the things you learn doing thing N transfer to thing N+1, so you get economies of scale.
- taink 4y agoHow do you plot graphs with org? I've been trying to use it for that purpose but I can't wrap my head around how to do it without some tikz incantation I don't really understand. I've seen gnuplot mentioned here and there but the setup seems pretty involved. I'm looking for a way to plot simple numeric data signals in time series, which are pretty trivial in jupyter notebooks.
- V1ndaar 4y ago
- davidpolberger 4y agoBack in 2010, I used plasTeX (http://plastex.github.io/plastex/ http://plastex.github.io/plastex/) to convert my thesis to HTML (http://www.polberger.se/components/ http://www.polberger.se/components/). plasTeX is "a Python package to convert LaTeX markup to DOM." If memory serves, plasTeX worked rather well, and still seems to be maintained today.
- DominikPeters 4y agoI would recommend using the lwarp package for turning large latex documents into HTML. Pretty much all other converters attempt to parse the tex files, which is an almost hopeless task. Lwarp has a different strategy: it redefines all macros to produce HTML (e.g. \textbf{example} writes "<strong>example</strong>" into the output pdf) within latex, thereby producing a PDF containing HTML code. It then uses a pdf2txt extractor to get the finished HTML file. Thus, it uses latex to parse the latex. Lwarp worked for me to produce an HTML version of the TikZ documentation (https://tikz.dev https://tikz.dev), and that's probably one of the more complicated tex documents that exists. (Though granted, this was still a major effort.)
- gdprrrr 4y agoYeah, it's well known that only LaTeX can parse LaTeX because you can redefine all syntax (catcodes) in the middle of the document.
- periheli0n 4y agoThe real shocker is that it’s 2022 and LaTeX is still the best writing environment for a PhD thesis. It has so many downsides: the markup syntax is ugly, it really works best only if one used paginated output such as PDF, a zoo of partly incompatible packages, need for compilation, obscure figure placing algorithms that are difficult to control, and so on. It still beats the competition because of rock-solid referencing, both to in-text elements like equations, chapters, etc as well as citing literature with bibtex. Plus, it’s extremely stable, so someone who learnt LaTeX 20 years ago, like yours truly, can download the newest TeX distribution and feel at home immediately. Nevertheless, I would prefer a Markdown-based system that can use CSS and MathML, and has a 100% bibtex clone for references. Yes, pandoc goes quite a long way along this route, but setting up such a pipeline is still too complicated for many.
- chaoxu 4y agoHave you tried Quarto? It should tick everything in your box (except MathML, but hey that might work too since Quarto is built on pandoc)
- periheli0n 4y agoThanks for the pointer, that looks interesting. Especially because it is open source! I see it supports Jupyter notebook. Math support in those isn’t too bad at all, so it might just work for many cases.
- countrymile 4y ago+1 for quarto, i wrote my thesis in rmarkdown which flipped easily between latex and html output, with a bibtex referencing system. It also allowed you to inline latex for more complex outputs. And inlining calculated tables and charts meant i could keep my writing and code together. Quarto is the successor.
- runningmike 4y agoI would strongly recommend MyST. MyST extends Markdown for technical and scientific communication. See https://www.myst.tools/ https://www.myst.tools/
- BrandoElFollito 4y agoIt was very cool for OP to do a writeup of the effort it took to convert a thesis. I wanted to do the same with mine but I lost the sources (it was 20+ years ago and did not survive some upgrade/technology change). I was mostly concerned with the .eps files that are hardly portable to .png or similar. This made me think a bit about preservation of ~recent data (1990-2010). a lot falls in the category of "not natively on the web yet" and "stored on stuff that does not work anymore".
- the-printer 4y agoIf he expects people (me) to read or read about his online PhD thesis then I think he should’ve chosen a font with a larger x-height. Reading the type feels like peeking through a dense bush during a hail storm.
- bspammer 4y agoThe fact that it's published as HTML means that you can choose your own font if you so desire. A PDF wouldn't let you do that.
- nicodjimenez 4y agoMathpix (https://mathpix.com https://mathpix.com) provides a drag and drop tool that converts PDF -> Markdown -> (HTML, LaTeX, PDF, DOCX). It's very handy and a lot of researchers and publishers use this tool, as well as people in the accessibility space (we make math content accessible for visually impaired students). Disclaimer: I'm the founder of Mathpix.
- kvakkefly 4y agoA great alternative to latex is doconce. It can output html, latex and various other formats. http://hplgit.github.io/doconce/doc/web/index.html http://hplgit.github.io/doconce/doc/web/index.html
- patrick451 4y ago> PrinPDFs are difficult to browse, impossible to read on a phone, uncomfortable to read on a tablet, hostile to screen readers, impractical to search engines, and the list goes on. It's just a terrible format, unless you're trying to print things on paper. ting things is a perfectly reasonable thing to do, but that's really not the main use case we should be optimizing for. The printed page is the optimal format for consuming a thesis or paper, so it absolutely makes sense to optimize for that, rather than dilettante browsing on a phone.
- easygenes 4y agoI have almost the opposite opinion to the author on the use of PDFs. I quite enjoy the experience of reading PDFs as opposed to just about any other e-document format, at least for the types of documents where PDFs are typically an option (e.g. books, magazines, and scholarly papers). In particular, the points they specifically call out about the difficulties with PDFs are either contrary to my experience or irrelevant to me: They cite: "... difficult to browse, impossible to read on a phone, uncomfortable to read on a tablet, hostile to screen readers, impractical to search engines ..." Browsability: I use Calibre [1] as an e-document library and PDFs are a first class citizen. It is a lovely piece of software for which I have few complaints and much praise. Phone and tablet readability: I've used a mix of GoodReader and more recently Documents on my iPhones and iPads and have never had troubles with reading PDFs on them. Screen reader access: I'm not much of a screen reader user, so no comment. Impractical to search: While you may occasionally need to go to an extra step to search PDF collections, it's hardly impractical and there are many performant options. Calibre includes a full-text indexing feature for your whole document library, to name one. Am I in the minority here for actually preferring PDFs? 1: https://calibre-ebook.com/