2 ms·
I've always taken particular note of how wisely scope-limited Pandoc is. Markup that aligns with natural language convention[1], is tightly converted, but the f
by lopsotronic 2mo ago
I've always taken particular note of how wisely scope-limited Pandoc is. Markup that aligns with natural language convention[1], is tightly converted, but the further from natural language, the less fidelity Pandoc can promise. Until, at the DITA or S1000D stage of "this ain't natlang, brah", Pandoc says "forget it" and just won't even pretend that such markup is even convertible.
Because, spoiler, it's not.
Constructs like tables and bibliographies challenge natural language markup - resulting in an explosion of different formalisms - but component content system artifacts for transclusion and conditionals shatter any pretense that these file types are "documents" at all. Both of those artifacts must draw formal structure from outside of language, i.e., from their own product / domain. They're parts of a system that make documents, but are not themselves documents or natural language. They are meaningless - or, worse, full of wrong meaning - outside of their runtime environment inside an explicit knowledge domain. Something that newer component content formats like Typst recognize explicitly.
The proof of all this is, as they say, in the pudding. What do people write documents in today? Well, they stick to natural language formats, sometimes they let the document system handle tables in some bespoke way, but conditionals are viewed with justified suspicion. DITA and S1000D projects, and the cursed migrations that lead to them, are sparse and driven almost exclusively by regulatory requirements, or, more often, program offices misreading regulatory requirements[0].
And here we all are in the LLM age, where natural language is being vindicated in ways both awe-inspiring and devastating. While component content systems force an LLM to expand its context window to the entire repository to make sense of any single sentence.
The crap of all this is, this is stuff that computer / information science has known since at least the 1980s. There are papers written about it. But high-complexity component content systems are sellable to non-technical writer groups because they don't see the tripwires in the fundamentals, or they think[2] that their product domain is so structured that the tripwires can be rigged as structure.
[0] No, converting to a pile of S1000D 040As doesn't magically fix your MTAs or your ILS or anything else.
[1] I do realize that proximity to natural language is correlate, not cause. Markdown converts well as a low-power notation whose instances denote values; it reads like natural language because that's what low-power does. The operative variable is whether the artifact denotes a document or a function from configuration to documents. AsciiDoc with `ifdef::[]` and `include::[]` converts every bit as badly as DITA, although without the fundamental nonsense of XSD and Horn's Information Mapping.
[2] Almost always wrongly