3 ms·
Responding to: EDIT: Looking at the spec [1] it seems to address some of these, but still indicates a strong confusion between data types (Unicode, rational nu
by bascule 10y ago
Responding to:
EDIT: Looking at the spec [1] it seems to address some of these, but still indicates a strong confusion between data types (Unicode, rational numeric) and data representations (UTF-8, IEEE double).
The format is described in terms of the tags (which act as type annotations), each of which corresponds to a specific on-the-wire format. Different tagged serializations of the same data may correspond to data of the same type. A better place to discuss ambiguities in the spec regarding this issue is here: https://github.com/tjson/tjson-spec/issues/27 https://github.com/tjson/tjson-spec/issues/27
The idea that different on-the-wire representations of an object correspond to the same typed data object (and can therefore result in the same hash) is core to understanding content-aware hashing.
So to your I'm guessing the author just conflated accusations, I don't think you fully understand what's going on here.
- colanderman 10y agoJSON is not defined in terms of UTF-8. That would be patently ridiculous, since UTF-8 is a serialization. JSON is defined in terms of Unicode code points. A string in JSON is a sequence of code points, some of which are (necessarily) escaped, others of which may be. So, to say "the string must be UTF-8" makes no sense. The JSON serialization itself can be UTF-8 (which I presume is what the author means). But nowhere does JSON talk about the encoding of a string within JSON, because it is not encoded. Furthermore, what does the author intend for escaped characters? Are they allowed? Presumably not, since that would provide for non-canonical representations. But some escapes must be allowed, since control characters (i.e. code points less than U+0020) must be escaped per the JSON spec. Nowhere does he address this; just a technically meaningless "strings must be UTF-8".
- bascule 10y agoJSON is not defined in terms of UTF-8. That would be patently ridiculous, since UTF-8 is a serialization. TJSON is defined as a serialization format on top of a JSON-like data model. The TJSON spec originally used the terminology "Unicode String", but moved to using "UTF-8 String", the rationale for which is given here: https://github.com/tjson/tjson-spec/issues/27 https://github.com/tjson/tjson-spec/issues/27 If your intent is to actually effect a change in the specification, that is the proper place to do it, but specific criticisms of the exact wording of the specification, preferably in the form of pull-requests, would be the best way to affect such changes. If your intent is not to effect a change in the specification, you're entitled to your opinion, but I'm done discussing the matter as the discussion has ceased to be meaningful to me. Generic criticisms like "You used 'UTF-8' instead of 'Unicode'" outside the context of specific sections of the specification aren't particularly helpful. Furthermore, what does the author intend for escaped characters? Are they allowed? Presumably not, since that would provide for non-canonical representations. You are continuing to miss the point: TJSON intends to provide a foundation to use content-aware content hashing in lieu of a canonicalization scheme as an alternative solution which works across multiple encodings of the same data, sidesteps the exact problems you're talking about, and also allows arbitrary subsets of an object graph to be authenticated without requiring rehashing/resigning. Please see this closed issue on canonicalization ("won't do"): https://github.com/tjson/tjson-spec/issues/24 https://github.com/tjson/tjson-spec/issues/24 From what I can gather, TJSON is offering a degree of abstraction you have not yet fully gleaned. The core idea is: many serializations, one underlying data structure/object graph. TJSON is a mere serialization layer, and indeed many TJSON documents may refer to the same underlying data structure, but all will have the same "objecthash": https://github.com/benlaurie/objecthash https://github.com/benlaurie/objecthash
- colanderman 10y ago> The core idea is: many serializations, one underlying data structure/object graph. Then why is UTF-8 even mentioned? Or time zone offsets, for that matter?
- bascule 10y agoSo it's possible to specify a rigorous set of tests cases that, ideally if all are passed, can be used to certify a conforming implementation. In other words, to solve this problem: http://seriot.ch/parsing_json.php http://seriot.ch/parsing_json.php While in some cases it might make sense to relax some of the requirements, I'm a fan of keeping things simple. Call me one of those crazy people who thinks Postel's Law is wrong. TJSON specifies a set of test cases for this purpose here: https://raw.githubusercontent.com/tjson/tjson-spec/master/draft-tjson-examples.txt https://raw.githubusercontent.com/tjson/tjson-spec/master/dr... I prefer to specify things in such a way that it's relatively easy to specify a test suite that covers all of the corner cases. A secondary goal of TJSON is to produce a stricter format, so I'd prefer to start with additional strictness requirements, and relax them if a reasonable case can be made.