3 ms·
JSON is not defined in terms of UTF-8. That would be patently ridiculous, since UTF-8 is a serialization. TJSON is defined as a serialization format on top of
by bascule 10y ago
JSON is not defined in terms of UTF-8. That would be patently ridiculous, since UTF-8 is a serialization.
TJSON is defined as a serialization format on top of a JSON-like data model. The TJSON spec originally used the terminology "Unicode String", but moved to using "UTF-8 String", the rationale for which is given here: https://github.com/tjson/tjson-spec/issues/27 https://github.com/tjson/tjson-spec/issues/27
If your intent is to actually effect a change in the specification, that is the proper place to do it, but specific criticisms of the exact wording of the specification, preferably in the form of pull-requests, would be the best way to affect such changes.
If your intent is not to effect a change in the specification, you're entitled to your opinion, but I'm done discussing the matter as the discussion has ceased to be meaningful to me. Generic criticisms like "You used 'UTF-8' instead of 'Unicode'" outside the context of specific sections of the specification aren't particularly helpful.
Furthermore, what does the author intend for escaped characters? Are they allowed? Presumably not, since that would provide for non-canonical representations.
You are continuing to miss the point: TJSON intends to provide a foundation to use content-aware content hashing in lieu of a canonicalization scheme as an alternative solution which works across multiple encodings of the same data, sidesteps the exact problems you're talking about, and also allows arbitrary subsets of an object graph to be authenticated without requiring rehashing/resigning. Please see this closed issue on canonicalization ("won't do"):
https://github.com/tjson/tjson-spec/issues/24 https://github.com/tjson/tjson-spec/issues/24
From what I can gather, TJSON is offering a degree of abstraction you have not yet fully gleaned. The core idea is: many serializations, one underlying data structure/object graph. TJSON is a mere serialization layer, and indeed many TJSON documents may refer to the same underlying data structure, but all will have the same "objecthash":
https://github.com/benlaurie/objecthash https://github.com/benlaurie/objecthash
- colanderman 10y ago> The core idea is: many serializations, one underlying data structure/object graph. Then why is UTF-8 even mentioned? Or time zone offsets, for that matter?
- bascule 10y agoSo it's possible to specify a rigorous set of tests cases that, ideally if all are passed, can be used to certify a conforming implementation. In other words, to solve this problem: http://seriot.ch/parsing_json.php http://seriot.ch/parsing_json.php While in some cases it might make sense to relax some of the requirements, I'm a fan of keeping things simple. Call me one of those crazy people who thinks Postel's Law is wrong. TJSON specifies a set of test cases for this purpose here: https://raw.githubusercontent.com/tjson/tjson-spec/master/draft-tjson-examples.txt https://raw.githubusercontent.com/tjson/tjson-spec/master/dr... I prefer to specify things in such a way that it's relatively easy to specify a test suite that covers all of the corner cases. A secondary goal of TJSON is to produce a stricter format, so I'd prefer to start with additional strictness requirements, and relax them if a reasonable case can be made.