4 ms·
I love your comment. Thanks for taking the time to look so deeply into this! I'll respond to the main points and then expand on the details later. I'm not sur
by djedr 4y ago
I love your comment. Thanks for taking the time to look so deeply into this!
I'll respond to the main points and then expand on the details later.
I'm not sure what the authoritative source on S-expressions is (or even if there is one, which to me is a problem and part of the value proposition here), so I'll take R^7RS[1] as a reference.
If you look at the formal definition there (chapter 7), it's significantly more complex than both Jevko and what you beautifully constructed here.
And indeed, you have just freestyled a simplifed version of S-expressions (impressive!), but this is not the real thing.
If you would keep refactoring it with the constraints I had in mind for Jevko, you'd eventually end up with Jevko.
> though of course the name-value pairing is lost there (because S-expressions lack it).
Indeed, and that's another part of the value proposition of Jevko. The grammar is designed purposefully to take advantage of natural syntactic name-value (prefix-subjevko) pairing tendencies.
Which brings me to the next point.
**
You are absolutely correct that the abbreviated grammar matches the same strings, but doesn't have the same structure.
*The correct grammar is the one in the specification*.
I have shown the condensed version of it to illustrate the point that Jevko is indeed extremely simple. The single line captures all the essential elements and, again, matches the same strings as the full grammar.
This is unlike the similar condensed grammar for S-expressions which you sketched out here. If you would continue, it would get significantly more complex before it matches the same strings.
The OCaml type definition you wrote down should do the job of capturing the structure, although I prefer to name the elements (which may be not-so-convenient, depending on the language, so it's fine).
**
Indeed, I also think that this is an extremely useful data model. Thanks for pointing out the similarity to e-mail messages and other references which I'd love to dig into (please send if you have any links or resources about these databases).
Thanks for all the pointers, I'll take them into account.
And thanks again for your time.
[0] this seems kinda official, but there is no single standard here: https://www.s-expressions.org/standards https://www.s-expressions.org/standards
[1] https://standards.scheme.org/official/r7rs.pdf https://standards.scheme.org/official/r7rs.pdf
- kragen 4y agoThank you for your thought-provoking explorations, for the unwarranted flattery, and for taking the time to respond! Clearly we share an interest in flexible simplicity. As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '. The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions." As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example. One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents: x [y] z a [b] c concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important. I don't know if you saw the last time this topic came up I linked to https://ogdl.org/ https://ogdl.org/, which seems pretty close to a minimal rose-tree notation. They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes). Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
- djedr 4y ago> Thank you for your thought-provoking explorations, for the unwarranted flattery, and for taking the time to respond! Clearly we share an interest in flexible simplicity. Indeed we do! :) > As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '. For anybody trying to parse your definition: ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20 it's missing `*()` around the right-hand side: ex = *(("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20) ;P Other than that it's a indeed lovely minimal definition. And it allows binary strings! Now if such a definition was accessible as some formal spec and implemented in various languages the way JSON is, I'd probably not be inclined to rolling my own. Or maybe I still would, because it's fun! ;) > The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions." Sure. So, would you agree that with proliferation of Lisp variants the term "S-expression" became rather vague? > As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example. Holy moly, this looks very nice! I like the style. Kudos. > One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents: x [y] z a [b] c > concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important. Yes, being closed under concatenation is a feature I was aiming for and it indeed does bring with it this issue. Just something to have in mind when devising formats. A simple solution here is to disallow having anything other than whitespace in the suffix of a Jevko with > 0 children. Then, if a format converts these labels to keys in a map, trimming leading and trailing whitespace, there is no problem. This is how I did it here: * https://github.com/jevko/easyjevko.js https://github.com/jevko/easyjevko.js > I don't know if you saw the last time this topic came up I linked to https://ogdl.org/ https://ogdl.org/, which seems pretty close to a minimal rose-tree notation. Yes, I've seen OGDL before. It's pretty nice. A similar one is https://treenotation.org/ https://treenotation.org/ I have experimented with indentation-based syntaxes myself, before settling on brackets. I have found them to be problematic, at least because: * For complex structures they become less compact. * A grammar that correctly captures significant indentation can't really be written in pure *BNF. The way OGDL does it is this: [12] space(n) ::= char_space*n ; where n is the equivalent number of spaces (can be 0) [13] block(n) ::= '\' (comment|break) (space(>n) string break)+ I had the same idea. Simple enough, but still. Brackets are simpler to formalize and implement and not harder to explain. > They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes). Yes, that might be interesting to look at, thanks for pointing it out. I have thought about this and came up with some ideas, but haven't decided on anything. I was thinking more along the lines of having the path DSL be simply implemented on top of Jevko, not as a completely separate grammar. > Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it. It's a nice description, but I think a little to detailed and technical to fit into a tagline. Maybe a little explanatory article with the diagram included. Would probably look something like this: https://github.com/jevko/writing/blob/main/2022-01-10-jevko-visual-parse-trees.md https://github.com/jevko/writing/blob/main/2022-01-10-jevko-... Although I'd gladly see your take on it. ;)
- aidenn0 4y agoAs someone who has been using lisp for over 20 years, kragen's definition matches with what I think of when the term "S-expression" is used. Of course it's so minimal that the standard syntax of any non-toy lisp will have extensions. FWIW the most minimal attempted standard of an s-expression would be rivest's basic s-expressions, which, as an extension of djb's netstrings can represent arbitrary binary values: <sexpr> :: <string> | <list> <string> :: <display>? <simple-string> ; <simple-string> :: <raw> ; <display> :: "[" <simple-string> "]" ; <raw> :: <decimal> ":" <bytes> ; <decimal> :: <decimal-digit>+ ; -- decimal numbers should have no unnecessary leading zeros <bytes> -- any string of bytes, of the indicated length <list> :: "(" <sexp>* ")" ; <decimal-digit> :: "0" | ... | "9" ;
- djedr 4y agoYes, S-expressions come in many flavors, some more minimal than others, some binary. They are all truly wonderful. The most wonderful to me are the simplest ones, and Jevko grows out of the same spirit as them. However, it does not attempt to be a new flavor of S-expressions and diverges in ways which to me are worth looking at. I hope it can appeal and be useful not only to minimalist syntax enthusiasts. BTW Some time ago I've been also experimenting with binary versions of Jevko, certainly with inspiration from both netstrings and Rivest's csexps: https://github.com/jevko/binary-experiments#asttolengthprefixed3 https://github.com/jevko/binary-experiments#asttolengthprefi... Since then I had some more ideas which I hope to get around to implementing at some point.