5 ms·
> More or less, what I want from markup is to convert a text string into a document tree: enum Element { Text(String), Node { tag: String,
by djedr 4y ago
> More or less, what I want from markup is to convert a text string into a document tree:
enum Element {
Text(String),
Node {
tag: String,
attributes: Map<String, String>
children: Vec<Element>,
}
}
fn parse_markup(input: &str) -> Element { ... }
> Markup language which nails this perfectly is HTML.
The reason HTML nails it perfectly is because this is modeled after HTML.
If I were to make up a markup language, I wouldn't follow that model.
In particular I would get rid of attributes which to me are a restricted kind of children with a specialized syntax.
This is both unnecessary and undesirable in many cases.
The major problem with attributes is the <String, String> mapping. Once something is defined to be an attribute, it cannot be sensibly extended without creating an unnecessary problem.
For example the `class` attribute in HTML looks like it was originally designed to hold a single class name. Then people realized that it would be desirable to have multiple classes per element.
So, to keep things backwards-compatible, the value of the attribute was extended to hold a space-separated list of classes instead. Essentially creating a little DSL inside of the attribute's value.
If `class` was instead a kind of child, initially limited to a single instance per element, extending to multiple instances in a backward-compatible manner would not require introducing the DSL. It would be natural. Just allow many `class` children.
A valid argument in HTML in favor of attributes is conciseness. But that's an artifact of the syntax.
We could make up a language with syntax for nodes as concise as HTML attributes, eliminating that argument.
> It feels like there’s a smaller, simpler language somewhere
Certainly.
I have experimented with many different designs for such a markup language on top of Jevko[0]. One interesting design that trivially maps to HTML looks like this:
h1 [Title]
p [
paragraph
]
p [
[paragraph with a ]
a [
href=[...]
[link]
]
]
Another one looks like this:
[h1][Title]
[p][paragraph]
[p][
paragraph with a
[href[...] a][link]
]
Both are extremely simple and minimal, but also extensible and lend themselves to writing by hand.
[0] https://news.ycombinator.com/item?id=33287620 https://news.ycombinator.com/item?id=33287620
- typon 4y agoYou just described Lisp :)
- djedr 4y agoSurely you mean S-expressions. They're great, but not as a markup language. What I show here is in fact even simpler and more flexible than S-exps[0]. [0] For some details and polemic see this thread: https://news.ycombinator.com/item?id=33334789 https://news.ycombinator.com/item?id=33334789 | TL;DR: Jevko is well-defined, basically just unicode text + escapeable brackets for making trees; it doesn't treat whitespace as a separator/atmosphere (particularly important in markup); it takes advantage of natural name-value pairing tendencies (like tag-children); and it's closed under concatenation by design
- deleted 4y ago[deleted]
- notriddle 4y ago> If `class` was instead a kind of child, initially limited to a single instance per element, extending to multiple instances in a backward-compatible manner would not require introducing the DSL. It would be natural. Just allow many `class` children. That's not the operative difference between children and attributes in HTML. In HTML, if an element is unsupported, its contents are shown, while its attributes are ignored. This means, if a browser saw something like this: <p> <class>literature</class> <class>english</class> Billions of years ago, the Universe was created. This made a lot of people very angry, and has been widely considered a bad idea. </p> In browsers that don't support, or even predate classes, you would want them to be ignored. This only works if they're attributes.
- djedr 4y agoI was talking about is making up a new markup language, which doesn't have to inherit the distinctions and behaviors of HTML. If, in this language, you wanted to have this feature in combination with what I proposed, you could simply mark the attribute-like children, to inform the engine that it should apply different defaults for them and unmarked children that it doesn't recognize. This is what I do in the first variant of my language: a [ href=[...] [...] ] the `=` appended to `href` marks it as an attribute. This certainly simplifies things when translating to HTML. But if we were not constrained by the legacy of HTML then perhaps bothering with this when designing a new language would be unnecessary. Maybe not having this default behavior does not matter in practice and you can have a simpler language without it. To determine whether that's the case it would help to answer: when is the feature you described useful in HTML? And also: when is this feature harmful?
- notriddle 4y agoIf you just want to have your language throw up an error whenever it sees an unrecognized element, then you’ll probably be able to simplify it a lot. It’s probably fine, since * as long as you use a build tool the errors will be seen by the author (who will know how to fix them) and not the reader * graceful degradation and format extensions are a crapshoot due to Hyrum’s Law [1] [1]: https://news.ycombinator.com/item?id=30726668 https://news.ycombinator.com/item?id=30726668
- goto11 4y agoThe element/attribute distinction is because HTML is a markup language, where the primary content is text and markup is used to add structure and metadata. You can eliminate attributes to get a simpler format, but it isn't really a markup language anymore, since the distinction between text and markup is removed. It is just a generic data structure, like JSON or s-expressions.
- chrismorgan 4y agoThe text/markup distinction is one I’ve been thinking a lot about recently as I’ve been both working on my own lightweight markup language, and implementing a component-based website using Django/Jinja/Nunjucks/whatever-style templates. There are too many places where you end up with attribute-like syntax, but want to pass markup, and it sometimes becomes unclear whether you’re working with a string or markup, yet the two are definitely different—in HTML serialisation, one needs escaping, the other doesn’t; or in something like DOM, one’s a string and the other’s a fragment. I’ve become progressively more and more convinced that these kinds of template languages are quite bad at writing this sort of thing. Take this example of unfortunate mixing: {% somecomponent attr="value", title="<strong>A:</strong> B" %} <strong>C:</strong> D {% endsomeecomponent %} Trouble is that you basically have at most one markup slot, but many things semantically want more than one. The closest you can generally get in such languages as these is the likes of this: {% somecomponent attr="value" %} <strong>A:</strong> B {% body %} <strong>C:</strong> D {% endsomecomponent %} This can be done in some such languages, but not others—and can’t be done in anything like XML. In XML, the solution would be extra nesting, which gets messy quickly: <somecomponent attr="value"> <title> <strong>A:</strong> B </title> <body> <strong>C:</strong> D </body> </somecomponent> In JSX I suppose you could do something like this, for better or for worse: <SomeComponent attr="value" title={<><strong>A:</strong> B</>}> <strong>C:</strong> D </SomeComponent> If you use a fragment/child nodes for what is logically a string value, you run the risk of people putting non-text in there; things like an href should clearly be strings, not text nodes—or else you have to decide what to do with <a><href>/<b>exam</b>ple</href>…</a>, and poorly-written DOM manipulators will surely occasionally assume the href element contains only one text node at times. But there’s then also the question of whether you want strings to be usable in text node context (probably the more pragmatic approach), or if you want to maintain a strong distinction between the two (probably the more correct approach). For my own language, in the more macroy part (when you’re extending it beyond the basic elements), I’m maintaining a distinction between strings and markup fragments, using different delimiters for the two, with "…" for strings and {…} for fragments. I’m also treating attributes and children equivalently, having all as possibly-named children. The basic example could be something like this in my current syntax: @somecomponent(attr: "value", title: {**A:** B}, { **C:** D })