4 ms·
> Consider the approach taken by XML 1.1 I evaluated xᴍʟ 1.0 and 1.1’s restrictions but consider the alternative, accepting anything but the empty string to be
by GeneThomas 2y ago
> Consider the approach taken by XML 1.1
I evaluated xᴍʟ 1.0 and 1.1’s restrictions but consider the alternative, accepting anything but the empty string to be simpler.
>A lot of English-speaking countries actually have `.` and `,` swapped:
Which?
>...should state that encoders should pick a long enough decimal representation to avoid any such issue.
.net for example has round trip encoding to achieve this.
> Which base64 standard in [1]?
As linked in the document, we refer to ʀꜰᴄ 4648 for Base64 using the simplest version which states that “Implementations MUST include appropriate pad characters at the end of encoded data unless the specification referring to this document explicitly states otherwise”.
> I have outlined my rationale for layering in other comments.
>> Graph is generally a bad thing to encode at this level
The ᴅᴏꜱ attach you are concern about is not applicable. Layering graphs on top of the encoding in a separate standard is not practical. As stated serialization requires a graph structure.
>It is not too hard to define a data model in prose rather than code.
I have done that.
>May have been true in the past, but it's no longer true since SIMD-based parsers. Also the very existence of escape sequence prevents the true non-destructive zero-copy (aka in-situ) parsing anyway. With such sequences, zero-copy/in-situ parsing has to be destructive to be performant and that can preclude some use cases. Allowing additional space characters is much easier than that.
SMID based parsers would almost certainly run faster with simplified byte handling rules. Practically, when say skipping whitespace, searching for two specific bytes is faster than the alternatives. Most strings do not have escape sequences so do not need to be copied, a scan for '\' is fast.
> `<<Name>...<<$>`
This is to be terse. An array type in an xᴍʟ like language is a leap forward. Additional significant characters would over complicate the standard.
- MzHN 2y ago>>A lot of English-speaking countries actually have `.` and `,` swapped: >Which? Here's a table https://en.wikipedia.org/wiki/Decimal_separator#Examples_of_use https://en.wikipedia.org/wiki/Decimal_separator#Examples_of_... EDIT: Ah, now that I re-read that, I see the "English-speaking" specifier.
- lifthrasiir 2y ago> Which? Oh, I just realized that I carelessly put "English-speaking" there. (AFAIK the exact reversal does exist in English, but is much rarer and not entirely domestic.) But that doesn't really justify the use of comma in numbers. > .net for example has round trip encoding to achieve this. That was never guaranteed to my knowledge, and there also seems an apparent difference in .NET Framework and .NET Core/Runtime versions according to some searches. Many enough standardized languages have no strong requirement as well (for example, see [1] for ECMAScript), so this should be clearly specified to be truly portable. [1] https://tc39.es/ecma262/multipage/abstract-operations.html#sec-roundmvresult https://tc39.es/ecma262/multipage/abstract-operations.html#s... > As linked in the document, we refer to ʀꜰᴄ 4648 for Base64 using the simplest version [...]. You are free to rephrase any variant of base64 standard into simple statements. (Ideally the standard itself should be also cited, though.) I was asking why that particular variant was used. > Layering graphs on top of the encoding in a separate standard is not practical. As stated serialization requires a graph structure. The separateness here is not binary. The Unicode standard (ISO/IEC 10646) for example has multiple standard annexes [2] that are synchronized with the core standard but published and developed separately. For all practical purposes they form a single unified standard, but you can (and almost likely should) ignore some irrelevant annexes. I'm arguing for a similar structure to be clear, do you think that's also a no-go? [2] https://www.unicode.org/reports/ https://www.unicode.org/reports/ > I have done that. The only explicit thing in your data model is the following definition: // Read as: "document" is either "object", "array" or "scalar". document = object | array | scalar // Read as: "object" is a distinct type "Object" with "name" and zero or more "field"s. (And so on) object = Object(name, field*) array = Array(name, item*) scalar = Scalar(name, value) This is not what's implied by following sections. To begin with, you have no explicit definition of `item` or `field`. Yes, it is kinda obvious that both should be `document` and also can be somehow inferred that `scalar` in the `item` place should be a nameless object, but not only such relations are not explicit but now we have another class of objects not mentioned in the first place! The exact nature of `value` is also very unclear. In my understanding, your data model is more like this: document = object | array | scalar // `...?` for zero or one copy of the preceding data. optional-name = Name(name)? optional-ref-id = RefId(ref-id)? optional-type = Type(type)? name = string // Can't be empty ref-id = string // Globally unique type = string // Interpretation up to clients object = Object(optional-name, optional-ref-id, optional-type, field*) field = ObjectField(name, optional-ref-id, optional-type, field*) | scalar array = Array(optional-name, item*) item = object | ScalarItem(optional-ref-id, field*) // Implicitly converted to object | ReferenceItem(reference-id) scalar = Scalar(name, value) | Reference(name, ref-id) value = string // Additional interpretation rules may exist after parsing For example, the following document: <Person> <Name=Bonnie> <Spouse#jack-smith> <Name=Jack> <$> <Doctor=@jack-smith> <$> ...is thought to parsed into the following underlying data: Object(Name("Person"), Scalar(Name("Name"), "Bonnie"), ObjectField(Name("Spouse"), RefId("jack-smith"), Scalar(Name("Name"), "Jack") ), Reference(Name("Doctor"), "jack-smith") ) This abstraction clearly demonstrates many important things coming up in implementations. For example, every `field` has to provide `name` because all possible branches (`ObjectField`, `Scalar` or `Reference`) have one---at least in my understanding. I shouldn't be figuring out such data model myself to be honest; it's your job to provide one.