5 ms·
One comparison is here: https://jevko.github.io/compactness.html https://jevko.github.io/compactness.html To see how it's syntactically simpler than S-exps not
by djedr 4y ago
One comparison is here: https://jevko.github.io/compactness.html https://jevko.github.io/compactness.html
To see how it's syntactically simpler than S-exps note that the grammar of Jevko can be condensed into one short line of ABNF:
Jevko = *("[" Jevko "]" / "`" ("`" / "[" / "]") / %x0-5a / %x5c / %x5e-5f / %x61-10ffff)
The grammar of S-exps on the other hand, I won't quote here, but I assure you it's much more complicated. How much depends on your flavor (Jevko is also simpler in this regard: there is only one flavor, clearly specified).
There is no (intended) ambiguity around whitespace in Jevko: whitespace does not occur explicitly in the grammar. Whitespace characters are just characters. This is the defining feature of the syntax.
For this reason Jevko is more low-level: if you want to treat whitespace in some special way, you have to do that yourself. Although for most use-cases this is very similar and simple, e.g. https://news.ycombinator.com/item?id=33334314 https://news.ycombinator.com/item?id=33334314
But the point is that you can also leave it as-is, e.g.: https://github.com/jevko/queryjevko.js https://github.com/jevko/queryjevko.js
or do something else -- it's up to your format.
- simplotek 4y ago> To see how it's syntactically simpler than S-exps (...) Has anyone ever complained that S-expr were too complex? Adding white noise as a tradeoff doesn't seem like a win.
- avmich 4y ago> Has anyone ever complained that S-expr were too complex? Ask parser writers, I guess? I had some similar devices, explicitly made so that the overall set of parsing rules would be clear, so the idea is likely not too new.
- simplotek 4y ago> Ask parser writers, I guess? I've written many parsers, including for s-expr following rivest's RFC. S-expr take about half a dozen states to pull. A packrat parser for s-expr can fit in a single page. You don't even have to scroll to see the whole implementation. What are you talking about? What exactly do you feel parser writers will tell you regarding s-expr?
- kragen 4y agoHmm, wouldn't S-exps in the same ABNF notation be ex = ("(" *ex ")" / *%x2a-10ffff) *%x0-20 We could write the data model corresponding to this grammar in OCaml as type ex = List of ex list | Atom of string If we add double-quoted strings with GW-BASIC/SQL-style quoting (and resolve the ambiguity greedily): ex = ("(" *ex ")" / *%x2a-10ffff / %x22 *(%x22 %x22 / %x0-21 / %x23-10ffff) %x22 ) *%x0-20 This corresponds to the data model type ex = List of ex list | String of string | Symbol of string This still seems both simpler and more expressive than the one-line grammar you give. Maybe I'm missing a subtlety of S-expressions here, and I haven't tried it, but I think this correctly parses your examples like (first-name"John"last-name"Smith"is-alive true age 27 address(street-address "21 2nd Street" city"New York"state"NY"postal-code"10021-3100") phone-numbers((type "office"number"212 555-1234") (type "home"number"646 555-4567")) children()spouse()) or ( :first-name "John" :last-name "Smith" ... ) though of course the name-value pairing is lost there (because S-expressions lack it). (By the way, if you want to attribute your JSON example for copyright reasons, you need to attribute it to its author or authors, not to the Wikipedia, which is just the site they posted it on.) — ⁂ — Maybe more importantly, though, I think your abbreviated Jevko grammar is wrong in an important way, though it describes the same set of strings as the full grammar. According to the abbreviated grammar, the Jevko examples lack most of the structure of the XML, JSON, and S-expression versions. It implies that a Jevko is an ordered sequence of (Unicode!) characters and nested Jevkos (indicated by []). We could express this data model as type jevko = atom list and atom = Char of char | Nest of jevko If so, then abbreviating your S-expression example (first-name "John" last-name "Smith") the closest Jevko equivalent is not, as you claim first name [John] last name [Smith] but rather (supposing the hyphens were just an unfortunate concession to S-expression syntax rather than actually desired) [first name][John][last name][Smith] We had to sacrifice the formatting white space because there's nowhere that Jevko (as specified above!) ignores it. Using this grammar, the S-expression equivalent of the Jevko first name [John] is rather ("f" "i" "r" "s" "t" " " "n" "a" "m" "e" " " ("J" "o" "h" "n")) If we instead use the full Jevko grammar Jevko = *Subjevko Suffix Subjevko = Prefix "[" Jevko "]" Prefix = Text Suffix = Text Text = *Symbol Symbol = Digraph / Character Digraph = "`" ("`" / "[" / "]") Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff or, I think, equivalently in this context: Jevko = *Subjevko Text Subjevko = Text "[" Jevko "]" Text = *(Digraph / Character) Digraph = "`" ("`" / "[" / "]") Character = %x0-5a / %x5c / %x5e-5f / %x61-10ffff then we do preserve the name-value structure you seem to be going for, which your above one-line version loses. And this allows us to write first name[John]last name[Smith] as in your compactness examples. I think this is a more useful level of abstraction, and it's more or less the level used by, for example, queryjevko.js's jevkoToJs, although that erroneously uses () instead of []. (Also, contrary to your assertion above that this is an example of "leaving [Jevko's data model] as-is", it forgets the order of the name-value pairs as well as I guess all but one of any duplicate set of fields with the same name and also the possibility that there could be both fields and a body.) Essentially at this level of structure a Jevko is a (possibly empty) set of name-value pairs followed by a plaintext body ("Suffix"). This is exactly like an email message, except that the values are themselves Jevkos. In OCaml we could write: type jevko = Jevko of (string * jevko) list * string For your example include[author]fields[articles[[title][body]]people[[name]]] this gives the representation Jevko ([("include", Jevko ([], "author")); ("fields", Jevko ([("articles", Jevko ([("", Jevko ([], "title")); ("", Jevko ([], "body"))], "")); ("people", Jevko ([("", Jevko ([], "name"))], ""))], ""))], "") although I notice that queryjevko handles it very differently. — ⁂ — Unlike the data model implied by your one-line grammar, I think this is an extremely useful data model. Email messages are the one and only structured data format that has remained compatible in active use and extension for over half a century; you can literally take internet email messages from 01972 and load them into a mail client today (most easily by putting them into a qmail-style maildir) and everything will just work. This is largely a result of the decentralized extensibility properties of name-value pairs: mail clients just ignore header names they don't understand, and they don't require header names that weren't originally present. This is also the basis of the extensibility of HTTP/1.0 and HTTP/1.1. It's also very similar to the data model of popular "semistructured" or "free-form" databases of the 01990s like askSam and Filemaker. Unlike RFC-822 and HTTP/1, but like Jevko, those systems support recursively nested data. askSam even used almost the same firstname[John] syntax. This data model does have a semantic mismatch with things like the rose-tree representation you describe at https://xtao.org/blog/rose.html https://xtao.org/blog/rose.html, since it associates the rose-tree label with the first branch and the empty-string label with subsequent branches. If your audience is people like me, I think it would probably be worthwhile for you to spend some time up front describing the intended semantics of a data model, as I've attempted above, rather than leaving people to infer it from the grammar. (Maybe OCaml is not a good way to explain it, though.) You might also want to specify that leading and trailing whitespace in prefixes is not significant, though it is in the suffix ("body"); this would enable people to format their name-value pairs readably without corrupting the data. As far as I can tell, this addendum wouldn't interfere with any of your existing uses for Jevko, though in some cases it would simplify their implementations. ______ Runnable OCaml code: type jevko = Jevko of (string * jevko) list * string (* XXX doesn't escape `[ `] `` *) let rec dump (Jevko(hdrs, body)) = hdr(hdrs) ^ body and hdr = function [] -> "" | (k, v) :: t -> k ^ "[" ^ dump(v) ^ "]" ^ hdr(t) let dict kvs = Jevko(kvs, "") let text s = Jevko ([], s) let v = dict ["include", text "author"; "fields", dict [ "articles", dict ["", text "title"; "", text "body"]; "people", dict ["", text "name"] ]] ;; print_endline(dump v)