8 ms·
- The flexibility in name is to support the plethora of programming languages. - Commas are for readability as per English. - Three digits are recommended for
by GeneThomas 2y ago
- The flexibility in name is to support the plethora of programming languages.
- Commas are for readability as per English.
- Three digits are recommended for the encoder.
- Textual data is decoded into binary information in software so setting expectations, e.g. supporting ∞, facilitates interoperability by reducing impedance mismatches.
- If the implementation uses ɪᴇᴇᴇ doubles the rounding shall be as expected.
- Unicode surrogates are disallowed.
- We are following the Base64 standard exactly.
- It is essential for data to be structured as a graph, it simply occurs. Serialization on most data formats has the required support for graphs awkwardly layered on top of the encoding.
- The C# implementation defines the data model; essentially objects, array and scalars with multiple parents for nodes. The Formats section is to, as stated, facilitate interoperability.
- Limiting whitespace to spaces and tabs does speed processing. The ᴜᴛꜰ-8 bytes can be left in the input buffer rather than copied. All significant bytes are ᴀꜱᴄɪɪ.
- There are no unmatched angle brackets (that was a typo).
- I know the reasoning behind XML.
- lifthrasiir 2y ago> The flexibility in name is to support the plethora of programming languages. As far as aware, there is no language that allows a tab in its identifier. I'm aware of some (uncommon) languages that allow a space in identifiers though. After all, the goal of language is firstly to set what's valid or not and secondly to give a meaning to valid one. The valid set should be as maximally different from the invalid set as possible, but allowing virtually invalid characters in names is against this goal. Consider the approach taken by XML 1.1 (not 1.0) if you need an idea. > Commas are for readability as per English. A lot of English-speaking countries actually have `.` and `,` swapped: 3,141.592 vs. 3.141,592 for example. Due to this ambiguity, a comma as a grouping separator is heavily discouraged. You can instead use a underscore `_` (very common in programming languages) or a space (more preferred in human texts) without such concerns. > Three digits are recommended for the encoder. Also, some English-speaking countries use different group sizes (notably India). I'm personally fine with three-digit groupings despite of that, but that should be clearly specified at the very least. > Textual data is decoded into binary information in software so setting expectations, e.g. supporting ∞, facilitates interoperability by reducing impedance mismatches. This procedure should have been made explicit. In fact, the single worst thing you can do in serialization formats is an unclear definition of data model. > If the implementation uses ɪᴇᴇᴇ doubles the rounding shall be as expected. There is nothing like "as expected" in IEEE 754. The current rounding mode is a part of the execution state, so leaving it unspecified risks a non-deterministic interpretation even in the same execution. Either you should specify some rounding mode (most likely round-to-even), or you should state that encoders should pick a long enough decimal representation to avoid any such issue. > We are following the Base64 standard exactly. Which base64 standard in [1]? The padding in base64 is vestigial anyway and serves no practical need by now, so there is no strong reason to use a particular standard with the required padding. It is much more important to decide what to do with incorrect padded bits. [1] https://en.wikipedia.org/wiki/Base64#Variants_summary_table https://en.wikipedia.org/wiki/Base64#Variants_summary_table > It is essential for data to be structured as a graph, it simply occurs. Serialization on most data formats has the required support for graphs awkwardly layered on top of the encoding. I don't question that graph structures often occur naturally and existing schemes are often awkward, but I think that's more of the lack of co-developed standards. I have outlined my rationale for layering in other comments. > The C# implementation defines the data model; essentially objects, array and scalars with multiple parents for nodes. The Formats section is to, as stated, facilitate interoperability. The data model should be abstract enough to be truly interoperable. JSON suffered a lot from having no defined data model to this day, even when there was a soft-of-reference implementation by Crockford. It is not too hard to define a data model in prose rather than code. > Limiting whitespace to spaces and tabs does speed processing. The ᴜᴛꜰ-8 bytes can be left in the input buffer rather than copied. All significant bytes are ᴀꜱᴄɪɪ. May have been true in the past, but it's no longer true since SIMD-based parsers. Also the very existence of escape sequence prevents the true non-destructive zero-copy (aka in-situ) parsing anyway. With such sequences, zero-copy/in-situ parsing has to be destructive to be performant and that can preclude some use cases. Allowing additional space characters is much easier than that. > There are no unmatched angle brackets (that was a typo). I meant to refer to `<<Name>...<<$>`. Multiple grouping characters can allow for simpler syntaxes.
- GeneThomas 2y ago> Consider the approach taken by XML 1.1 I evaluated xᴍʟ 1.0 and 1.1’s restrictions but consider the alternative, accepting anything but the empty string to be simpler. >A lot of English-speaking countries actually have `.` and `,` swapped: Which? >...should state that encoders should pick a long enough decimal representation to avoid any such issue. .net for example has round trip encoding to achieve this. > Which base64 standard in [1]? As linked in the document, we refer to ʀꜰᴄ 4648 for Base64 using the simplest version which states that “Implementations MUST include appropriate pad characters at the end of encoded data unless the specification referring to this document explicitly states otherwise”. > I have outlined my rationale for layering in other comments. >> Graph is generally a bad thing to encode at this level The ᴅᴏꜱ attach you are concern about is not applicable. Layering graphs on top of the encoding in a separate standard is not practical. As stated serialization requires a graph structure. >It is not too hard to define a data model in prose rather than code. I have done that. >May have been true in the past, but it's no longer true since SIMD-based parsers. Also the very existence of escape sequence prevents the true non-destructive zero-copy (aka in-situ) parsing anyway. With such sequences, zero-copy/in-situ parsing has to be destructive to be performant and that can preclude some use cases. Allowing additional space characters is much easier than that. SMID based parsers would almost certainly run faster with simplified byte handling rules. Practically, when say skipping whitespace, searching for two specific bytes is faster than the alternatives. Most strings do not have escape sequences so do not need to be copied, a scan for '\' is fast. > `<<Name>...<<$>` This is to be terse. An array type in an xᴍʟ like language is a leap forward. Additional significant characters would over complicate the standard.
- MzHN 2y ago>>A lot of English-speaking countries actually have `.` and `,` swapped: >Which? Here's a table https://en.wikipedia.org/wiki/Decimal_separator#Examples_of_use https://en.wikipedia.org/wiki/Decimal_separator#Examples_of_... EDIT: Ah, now that I re-read that, I see the "English-speaking" specifier.
- lifthrasiir 2y ago> Which? Oh, I just realized that I carelessly put "English-speaking" there. (AFAIK the exact reversal does exist in English, but is much rarer and not entirely domestic.) But that doesn't really justify the use of comma in numbers. > .net for example has round trip encoding to achieve this. That was never guaranteed to my knowledge, and there also seems an apparent difference in .NET Framework and .NET Core/Runtime versions according to some searches. Many enough standardized languages have no strong requirement as well (for example, see [1] for ECMAScript), so this should be clearly specified to be truly portable. [1] https://tc39.es/ecma262/multipage/abstract-operations.html#sec-roundmvresult https://tc39.es/ecma262/multipage/abstract-operations.html#s... > As linked in the document, we refer to ʀꜰᴄ 4648 for Base64 using the simplest version [...]. You are free to rephrase any variant of base64 standard into simple statements. (Ideally the standard itself should be also cited, though.) I was asking why that particular variant was used. > Layering graphs on top of the encoding in a separate standard is not practical. As stated serialization requires a graph structure. The separateness here is not binary. The Unicode standard (ISO/IEC 10646) for example has multiple standard annexes [2] that are synchronized with the core standard but published and developed separately. For all practical purposes they form a single unified standard, but you can (and almost likely should) ignore some irrelevant annexes. I'm arguing for a similar structure to be clear, do you think that's also a no-go? [2] https://www.unicode.org/reports/ https://www.unicode.org/reports/ > I have done that. The only explicit thing in your data model is the following definition: // Read as: "document" is either "object", "array" or "scalar". document = object | array | scalar // Read as: "object" is a distinct type "Object" with "name" and zero or more "field"s. (And so on) object = Object(name, field*) array = Array(name, item*) scalar = Scalar(name, value) This is not what's implied by following sections. To begin with, you have no explicit definition of `item` or `field`. Yes, it is kinda obvious that both should be `document` and also can be somehow inferred that `scalar` in the `item` place should be a nameless object, but not only such relations are not explicit but now we have another class of objects not mentioned in the first place! The exact nature of `value` is also very unclear. In my understanding, your data model is more like this: document = object | array | scalar // `...?` for zero or one copy of the preceding data. optional-name = Name(name)? optional-ref-id = RefId(ref-id)? optional-type = Type(type)? name = string // Can't be empty ref-id = string // Globally unique type = string // Interpretation up to clients object = Object(optional-name, optional-ref-id, optional-type, field*) field = ObjectField(name, optional-ref-id, optional-type, field*) | scalar array = Array(optional-name, item*) item = object | ScalarItem(optional-ref-id, field*) // Implicitly converted to object | ReferenceItem(reference-id) scalar = Scalar(name, value) | Reference(name, ref-id) value = string // Additional interpretation rules may exist after parsing For example, the following document: <Person> <Name=Bonnie> <Spouse#jack-smith> <Name=Jack> <$> <Doctor=@jack-smith> <$> ...is thought to parsed into the following underlying data: Object(Name("Person"), Scalar(Name("Name"), "Bonnie"), ObjectField(Name("Spouse"), RefId("jack-smith"), Scalar(Name("Name"), "Jack") ), Reference(Name("Doctor"), "jack-smith") ) This abstraction clearly demonstrates many important things coming up in implementations. For example, every `field` has to provide `name` because all possible branches (`ObjectField`, `Scalar` or `Reference`) have one---at least in my understanding. I shouldn't be figuring out such data model myself to be honest; it's your job to provide one.