4 ms·
> does pretty poorly at importing LaTeX As far as I know, an importer that interprets arbitrary LaTeX cannot exist (i.e. you have to run LaTeX on the file) bec
by GiovanniP 5y ago
> does pretty poorly at importing LaTeX
As far as I know, an importer that interprets arbitrary LaTeX cannot exist (i.e. you have to run LaTeX on the file) because the grammar of LaTeX documents is Turing-complete. There is a StackExchange Q&A where this is explained, but I do not understand the arguments ;-)
- quietbritishjim 5y agoThat makes sense. Also (and I realise this is a bit less deep than undecidable grammar) I've certainly seen LaTeX documents with so many random macros that I don't recognise the content of it at all! I'm sure we've all come across those. I can understand why a program wouldn't stand much of a chance. Personally I never really needed LaTeX import so I don't really see it as a problem, but I know it matters to some people.
- GiovanniP 5y agoOn the "import from LaTeX" Joris van der Hoeven advanced the argument that the impossibility of writing a general import file from LaTeX allows LaTeX to maintain its status as the main document preparation system for scientific articles. In fact, if you could import arbitrary LaTeX, an user of a different system could collaborate with LaTeX users (probably a form of "conservative conversion" would be necessary for back and forth translation between document formats).
- ChrisLomont 5y agoLaTeX imports and interprets arbitrary LaTeX, so other programs can theoretically do just as well. Importers running up against Halting Problems occur in many document formats, but this is not an issue in practice since most documents are not nefarious, and any importer can (and should) have sensible timeouts for parsing.
- mgubi 5y agoNo, document formats are not turing complete. HTML for example can be parsed without any problem. However take this TeX snippet "\foo [bar]" and tell me if "bar" is an argument of "foo" or not, without running the macro (which maybe you do not have. This is the basic problem any TeX conversion program encounters. TeX is not a document format but a programming language. The only thing you can do with a TeX program is to run it. On the other hand an HTML document can be parsed without knowing anything about its content.
- cardiffspaceman 5y agoI once upgraded an EPS interpreter that was used as the key piece of an EPS importer. I added a couple of PS things like 'arcto'. It seems to me a LaTeX importer could work as an interpreter, too. Probably not very much fun after a while.
- mgubi 5y agoGraphics is different since usually the graphics primitives are the common substrate of all graphics languages, so the interpreter can just be adapted to a different backend. For more structured content this is not possible: the natural output of TeX is a sequence of graphics boxes for the glyphs, or something like a DVI format. It is too late to retain semantic informations, e.g. about mathematical formula. Ask anybody which writes LaTeX/HTML converters for example. TeXmacs does a good job to interpret "usual" LaTeX files but you cannot expect to be able to interpret arbitrary files and still retain semantic informations, e.g.: sectioning, references, emphasised texts, theorem statements, change of languages (e.g. from English to C++ to Math...) etc...
- GiovanniP 5y ago> For more structured content this is not possible: the natural output of TeX is a sequence of graphics boxes for the glyphs, or something like a DVI format. It is too late to retain semantic informations, e.g. about mathematical formula. I think this applies to the natural output of TeX's "mouth" as well, i.e. one could exapnd LaTeX macros till one obtains TeX primitives, but at that point the information on structure is already lost.
- quietbritishjim 5y ago> LaTeX imports and interprets arbitrary LaTeX, so other programs can theoretically do just as well. Right, but the output is a typeset document. The fact that LaTeX (the program) can do this is no evidence at all that it would be possible to build a program that converts LaTeX (the format) into some intermediate representation that make sense to a human. For example, assuming that macros are recursively expanded in place, it's possible that there's no clear dividing line between "this is a user macro that should be expanded before exporting" and "this is an internal macro that should not be expanded" (e.g. you probably want to leave "theorem environment" from AMS package unexpanded because that is semantic information you'd want when you edit the document). As another example, what should be a table or equation etc. might end up as individual lines and characters floating around as separate elements on a page.