3 ms·
> Thank you for your thought-provoking explorations, for the unwarranted flattery, and for taking the time to respond! Clearly we share an interest in flexible
by djedr 4y ago
> Thank you for your thought-provoking explorations, for the unwarranted flattery, and for taking the time to respond! Clearly we share an interest in flexible simplicity.
Indeed we do! :)
> As for S-expressions, I agree that what I wrote above is a pretty minimal definition of S-expressions, but I still think it's valid as a definition of "S-expression", even though you'd use a modified definition in some environments. In particular, I think it successfully analyzes the structure of the S-expressions you use as examples. And I agree that it doesn't match the same strings as those modified versions; for example, it doesn't recognize $ as a symbol, or allow '.
For anybody trying to parse your definition:
ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20
it's missing `*()` around the right-hand side:
ex = *(("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20)
;P
Other than that it's a indeed lovely minimal definition. And it allows binary strings!
Now if such a definition was accessible as some formal spec and implemented in various languages the way JSON is, I'd probably not be inclined to rolling my own. Or maybe I still would, because it's fun! ;)
> The lexical syntax specified in R⁷RS §7.1.2 is, I think, a lot more complex than S-expressions, and it does not purport to define S-expressions; it includes Unicode, booleans, bytevectors, strings, piped symbols, vectors, circular references, quote, quasiquote, unquote, unquote-splicing, hexadecimal numbers, octal numbers, binary numbers, floating-point numbers, and so on. LISP 1.5 couldn't handle any of these, but I would be very uneasy with the assertion, "LISP 1.5 couldn't read S-expressions."
Sure. So, would you agree that with proliferation of Lisp variants the term "S-expression" became rather vague?
> As an example of a more fleshed-out S-expression language, a few years ago when I wrote Ur-Scheme (a Scheme-subset compiler in the language it compiles), my S-expression parser supported lists (with optional dotted tails), quote, string literals, symbols, comments, signed decimal integers, booleans, and character literals. (Check out the section "Actual Parsing" in http://canonical.org/~kragen/sw/urscheme/compiler.scm.html http://canonical.org/~kragen/sw/urscheme/compiler.scm.html.) This is evidently enough to conveniently write a compiler in, and I think some of it doesn't really belong to S-expressions per se — quote, comments, and booleans, for example.
Holy moly, this looks very nice! I like the style. Kudos.
> One problem of my suggested data model is that, although the language is closed under concatenation (unlike, say, HTML, XML, JSON, or RFC-822), that concatenation is not semantics-preserving. Given these two Jevko documents:
x [y] z
a [b] c
> concatenating them changes the label for [b] from "a" to "z\na", and perhaps more damningly, erases the whitespace before "z". But, since none of the alternative formats (except ndjson and I guess plain uninterpreted binary, ASCII, or Unicode) is closed under concatenation, maybe that's less important.
Yes, being closed under concatenation is a feature I was aiming for and it indeed does bring with it this issue.
Just something to have in mind when devising formats. A simple solution here is to disallow having anything other than whitespace in the suffix of a Jevko with > 0 children. Then, if a format converts these labels to keys in a map, trimming leading and trailing whitespace, there is no problem. This is how I did it here:
* https://github.com/jevko/easyjevko.js https://github.com/jevko/easyjevko.js
> I don't know if you saw the last time this topic came up I linked to https://ogdl.org/ https://ogdl.org/, which seems pretty close to a minimal rose-tree notation.
Yes, I've seen OGDL before. It's pretty nice. A similar one is https://treenotation.org/ https://treenotation.org/
I have experimented with indentation-based syntaxes myself, before settling on brackets.
I have found them to be problematic, at least because:
* For complex structures they become less compact.
* A grammar that correctly captures significant indentation can't really be written in pure *BNF. The way OGDL does it is this:
[12] space(n) ::= char_space*n ; where n is the equivalent number of spaces (can be 0)
[13] block(n) ::= '\' (comment|break) (space(>n) string break)+
I had the same idea. Simple enough, but still. Brackets are simpler to formalize and implement and not harder to explain.
> They also put some thought into designing a query language for their rose-tree-like data model, which might be adaptable to Jevko — though they label only nodes, and Jevko labels both nodes (with suffixes) and arcs (with prefixes).
Yes, that might be interesting to look at, thanks for pointing it out. I have thought about this and came up with some ideas, but haven't decided on anything. I was thinking more along the lines of having the path DSL be simply implemented on top of Jevko, not as a completely separate grammar.
> Maybe that's the subtitle for Jevko? "A minimal Unicode syntax for ordered trees with labeled nodes and labeled arcs." If that's the intended semantics it would be pretty easy to whip up a diagram in Dot to illustrate it.
It's a nice description, but I think a little to detailed and technical to fit into a tagline. Maybe a little explanatory article with the diagram included. Would probably look something like this:
https://github.com/jevko/writing/blob/main/2022-01-10-jevko-visual-parse-trees.md https://github.com/jevko/writing/blob/main/2022-01-10-jevko-...
Although I'd gladly see your take on it. ;)
- kragen 4y agoI don't think your revision of the "S-expression" grammar is in accordance with the normal use of the term; it says ()() is an S-expression, when conventionally it's considered to be two of them. If that's what you want, then adding the outer *() is fine, but you should probably drop the leading * on the recursive call *ex, because otherwise the grammar is ambiguous. If you want to allow binary strings you'd probably want to define the language over bytes rather than Unicode code points, though Markus Kuhn's UTF-8B (implemented in Python as "surrogateescape") can give you the best of both worlds. Yes, I agree that "S-expression" is rather vague, much like "CSV". If you disallow having anything other than whitespace in a Jevko with >0 children, Jevko becomes a rose-tree notation, except that the root is unlabeled, so really it's more like a rose-forest notation. If you want a rose-forest notation, you can get it in a less irregular way by declaring that Jevko infers an extra [] following a suffix containing any non-whitespace character, so foo[bar] is equivalent to foo[bar[]], foo[bar[]baz] is equivalent to foo[bar[]baz[]], and a[b]c is equivalent to a[b[]]c[], but foo[bar[] ] is equivalent to foo[bar[]] and not foo[bar[][]]. This change would foreclose the possibility of having significant leading or trailing whitespace in some places but not others, the way I was suggesting. Rose trees or rose forests are definitely simpler than Jevko's current data model, and they are equally powerful. Agreed about indentation. With respect to diagrams, git clone http://canonical.org/~kragen/sw/pavnotes2.git http://canonical.org/~kragen/sw/pavnotes2.git and look at {horse,johnsmith,player}.{jpg,jevko,dot}. For the diagrams I've taken the liberty of ignoring leading and trailing whitespace on arc labels (prefixes), as well as suffixes consisting only of whitespace. johnsmith and horse are your examples, but they just use Jevko as a terser syntax for JSON that suffers a whitespace problem. player.jevko is an example I cooked up based on Minetest's relational database schema, recast as a hierarchical schema. It also takes advantage of the ordered nature of subJevkos, the possibility of multiple identical prefixes in the same parent node, and the possibility of having both "headers" (subJevkos) and a "body" (suffix) in the same node. (As far as I’m concerned, everyone is free to redistribute these nine files, in whole or in part, modified or unmodified, with or without credit; I waive all rights associated with them to the maximum extent possible under applicable law. Where applicable, I abandon their copyright to the public domain. To the extent that I wrote them at all, I wrote and published them in Argentina in 2022. Today, in fact. But you wrote horse.jevko and johnsmith.jevko, deriving them from things on Wikipedia, so my abandonment of copyright here shouldn't be construed as claiming that I wrote them.)