6 ms·
Sorry but I would never use this format for both manual or programmatic approach. * I've tried to read the data this format describes without reading its docum
by tripple6 2y ago
Sorry but I would never use this format for both manual or programmatic approach.
* I've tried to read the data this format describes without reading its documentation and I just failed: the format is amazingly counter-intuitive. I never had a readability and understanding issues with XML/HTML, JSON or even YAML (that I think is overly complicated) when I saw them for the first time.
* Terse does not mean cryptic. Basic notation is just weird: why would it need unbalanced the less-than symbol to open the array? Why `<&>` for delimiting elements? Why `<<$>` but not `<$>>` at least just to be more readable by human and look balanced? The syntax goes more weird for arrays containing objects: indents (okay to some extent), `<>` and `<&>` (`{` and `}`?).
* Auto-removing whitespaces may hurt. If the format offers this, would it also offer a heredoc-style text like `cat <<EOF` in Bash so that the formatting could be preserved as is? `xml:space` and JSON string literals were designed exactly for this. (upd: I just saw new symbol: `|`... Well, okay, but another special character now.)
* Native support for arrays. I mentioned a few above. `<<Faults$$>` and `<<$$>` -- guess what these two mean if you see this first time? You would never guess. It's an empty array and an empty element, you've just failed.
* Graphs... Another weird syntax comes into the room: `#id;` but `@id` (no semicolon?). Okay, these seem to be first-class ids and refs, not necessarily designed for graphs (I'm not sure if the `#ID;` and `@` would play perfect with any non-empty names.) But what does graphs make first-class citizens here and why? Graphs can be expressed, I believe, in any data/markup format/language and then processed with a particular application if graphs are needed. By the way, arrays and objects are not necessarily trees from the semantic point of view. More graph processing issues were mentioned in other comments to this topic. What about the first-class support for sets? I'm kidding
* Comments. Another symbol here to come: `%`. To be honest, I can't recall any instance I could see the percent sign elsewhere for this purpose. What if the comments would start with a well-known `#` at least with a space right after it so that it wouldn't be considered a "graph id" (or, don't get me wrong, with another `<`/`$` sequence)
* Just got to the Escaping section and now I see how the characters are escaped. Perhaps this is okay.
* Scalars. Crazy number formatting and locale issues are waiting. The never-on-keyboard infinity symbol would be great for APL, but why not just Inf(inity)? Whatever the scalar value is, no need to cover all existing primitive scalars -- just let them be processed by an application since all scalars are text semantically. Another crazy things: what does make UUIDs that special for this format?; why does make Base64 that special so that it has native support (would it support Base16 for human-readable message digests; or Base58 to remove visually lookalike Base64 characters)?
* CR/LF? I can understand its semantic purpose, but why not LF to make it even more "blazingly" fast? Say good-bye to UNIX users.
* The cognitive load for the markup syntax absolutely does not make it efficient in typing. Believe me, it does not.
What I would do, I would probably enhance the widely used formats, say make JSON, which I find almost perfect from the syntax point of view, not require quotes for object property names if the names would not contain special characters like `:` just like it goes in JavaScript. And perhaps make XML "v2" move away from SGML hence loosening its syntax to get rid of the closing tags with shorter notation, first-class array support and fixing syntax issues especially for CDATA and comments that can't support `--`. You would blame me, but I love XML the most: it just has the richest set of standardized amazing well-designed extensions to operate XML with regardless the heavy XML syntax.
P.S. How does it look like in the document it marks up is minified (e.g., no whitespaces)?
- zzo38computer 2y ago> Comments. Another symbol here to come: `%`. To be honest, I can't recall any instance I could see the percent sign elsewhere for this purpose. PostScript is one programming language that uses a percentage sign for comments. TeX and METAFONT also use a percentage sign for comments. There are others, too.
- GeneThomas 2y ago> `<<$>` but not `<$>>` Having both begin and terminate arrays start with << is more consistent. > `<>` and `<&>` (`{` and `}`?). Using `{` and `}` would lead to more special characters. > Auto-removing whitespaces may hurt. It does not. > Graphs... But what does graphs make first-class citizens here and why? It is simpler to support graphs in the markup. The fact is that the data being serialized may be structured in a graph. > CR/LF It supports LF only ᴜɴɪx line ends as well as CR/LF internet line endings. > Comments [...] To be honest, I can't recall any instance I could see the percent sign elsewhere for this purpose LaTex and PostScript both use % for comments. # matches the usage in ᴄꜱꜱ and ʜᴛᴍʟ, relating to an id/page location. > What if the comments would start with a well-known `#` at least with a space right after it so that it wouldn't be considered a "graph id" Having a space after the # differentiate between and id and comment would be a mistake. > Scalars. Crazy [...] UUIDs The Formats section is to facilitate interoperability between implementations, e.g. if you are encoding a ɢᴜɪᴅ [easy to say] then format it this way. > not make it efficient in typing. It is more terse than ᴊꜱᴏɴ. > XML "v2" ... first-class array support Xᴇɴᴏɴ has first class array support, the xᴍʟ like syntax leads to the <empty-arrray$$> notation. > P.S. How does it look like in the document it marks up is minified (e.g., no whitespaces)? Good.
- tripple6 2y agoI love this. > Having both begin and terminate arrays start with << is more consistent. It hides context for humans. I am a human and I love to see what opens and what closes the context. Why would `<` open an array if `[` is astonishingly wide-spread practice? Why would `<<` close it just because you think it is more consistent? What if open/close balance is also consistency, especially for nested arrays? Also just think how many key strokes you'd save if you'd use `]` instead of [Shift]+`,` [Shift]+`,` [Shift]+`4` [Shift]+`.` if you declare it as readable text. > Using `{` and `}` would lead to more special characters. Agree. Too many now. > It is simpler to support graphs in the markup. The fact is that the data being serialized may be structured in a graph. I can't understand why you call it native graph support. The only thing it does is declaring an identified element and references to the element. I can't see how different is that comparing to XML or JSON that semantically "have graph support" just because they also can declare something considered ids and references to the identified element. > LaTex and PostScript both use % for comments. Yes, just learnt that from your comment and https://news.ycombinator.com/item?id=42047634 https://news.ycombinator.com/item?id=42047634 by zzo38computer. Thank you. > # matches the usage in ᴄꜱꜱ and ʜᴛᴍʟ, relating to an id/page location. No. The # symbol is overloaded: it may be a comment start, especially for line-oriented and human-readable text formats or scripts; CSS uses it for IDs; HTML has nothing to do with it since browsers only use # as a part of a URL to reference a particular identified element for navigation purposes only (it's called anchor in URL syntax; formerly web-browsers used <a name="anchor"> to navigate to a part of the page; as of now in the HTML5 world any `id` attribute is considered an anchor which I find a design flaw since ids are something to be used to identify hence any id from the document is exposed for navigation navigation purposes, but <a name="anchor"> is semantically something for navigation). > Having a space after the # differentiate between and id and comment would be a mistake. Of course it would in its current perspective if the id declaration is `#`. Don't know what `#<NON_WHITESPACE_CHAR>` would do if it's legal. > The Formats section is to facilitate interoperability between implementations, e.g. if you are encoding a ɢᴜɪᴅ [easy to say] then format it this way. I agree that it may look better for consistency purposes, but what interoperability is all that about? Why would formatting even affect it? From the consumer application point of view, it must be handled from its context defined by its purpose and semantic type. If my element/attribute is formally declared as a GUID, then why would I care that much if it's conventionally formatted? Would it be still a GUID if I encode it using Base64? The dashes in GUIDs are for humans only and they are optional, and the application knows it's a GUID to process it even leniently if it can. The same goes for ISBN/ISSN for books and magazines, card numbers, phone numbers, etc -- none of them require dashes or spaces or parentheses to be processed. This is why "Real numbers *should be stored* with commas for readability." is just hilarious. Why should? May I use underscores or dots or spaces to group digits (seriously, why comma)? Can I group digits after the period? If I need integers, why are they also limited to 32 bits and 64 bits? How would I present an arbitrary precision integer or non-integer number (say, I want the Pi number 197 digits after the 3)? If ∞ is allowed, but no mention on +Inf and -Inf, can be 4.2957×10^24 used instead of 4.2957e24? May I just have simple `D+(\.D+)?` for everything I need for true interoperability? I agree consistent formatting is really beautiful, but it must never be the key to process data. > It is more terse than ᴊꜱᴏɴ. Sorry, it's not. > Good Could you please provide an example of minified (a single line, no new lines) array of timestamps from your page? UPD: I've just seen https://news.ycombinator.com/item?id=42038508 https://news.ycombinator.com/item?id=42038508 by Oras . Well, you know. ---- In short, too many whys, weird syntax and design decisions, so I cannot see anything that makes it a "better alternative" to XML, JSON, or YAML.