5 ms·
Author here. I originally started this project because I wanted a consistent way to serialize JSON so that the serialized bytes would hash the same way every ti
by seagreen 10y ago
Author here. I originally started this project because I wanted a consistent way to serialize JSON so that the serialized bytes would hash the same way every time.
As I worked on it though I realized it might be of general interested to people. Thus the example in the README of piping JSON through multiple tools without generating trivial changes that mess up diffs.
Most of the decisions I made were clear: no insignificant whitespace, object keys must be ordered, etc.
There are two things I'm still not sure about:
+ Son doesn't provide escape sequences for any Unicode character that JSON allows to be written unescaped. This includes U+007f (ASCII "delete"). Will that cause a problem for many programs? All the other ASCII control characters are required to be escaped by JSON, U+007f is the only one left out.
+ Son doesn't allow trailing zeros in fractions. This means you can't serialize `1.0`, you have to serialize it as `1`.
I was confident in the decision to take out scientific notation (it would be cool if JSON parsers actually treated numbers as being in scientific notation and tracked significant digits, but they don't so I feel like that ship has sailed). Trailing zeros are different though because some JSON generators do use them to distinguish integers from fractions. The problem is that many parsers don't care about them, so you end up in a situation where parsers are tossing out information about documents, meaning they can't serialize them faithfully again which is the whole point of Son.
- mst 10y agoI am tempted to suggest either: 1. Keep 1.0 as a special case to maintain the int/float distinction (it's a float; calling it a fraction is kinda-of-a-lie). 2. Refuse to handle floats at all, at which point people can pass [ <mantissa>, <exponent> ] for reals or [ <numerator>, <denominator> ] for rationals. There is of course (3), "build a compliance suite and claim the parsers that toss out information are Incorrect", but that doesn't seem compatible with your postel-ish goals.
- seagreen 10y ago> it's a float; calling it a fraction is kinda-of-a-lie This happens not to be correct. By specification JSON numbers are just a series of characters, arranged in a certain way: https://tools.ietf.org/html/rfc7159#section-6 https://tools.ietf.org/html/rfc7159#section-6 In practice though many JSON parsers will parse non-integer numbers to floats. > 2. Refuse to handle floats at all, at which point people can pass [ <mantissa>, <exponent> ] for reals or [ <numerator>, <denominator> ] for rationals. This is a really interesting idea. If you're writing something that's super important like medical software it would probably be worth considering. However, my goals are just to make minimal changes that improve JSON some while still keeping it fairly readable, so I think that means I should stick with allowing `123.456` or whatever. I'd like to try to keep an open mind on this though.
- mst 10y agoOoh, my apologies wrt 'kind-of-a-lie'. I think all the parsers I've used inflate 1 to an int and 1.0 to a float ... that or they don't, and I just thought they did. Either way, many thanks for the correction.
- throwawaysed 10y agoinstead of messing with JSON and making it less human friendly (arguably the thing that made JSON so popular in the first place), why not just use something designed for efficient machine-to-machine transfer? Basically making JSON more machine friendly and stricter to parse undermines the original reason to use JSON. If efficiency and absolute correctness are important use something not designed to be forgiving to humans :)
- derefr 10y agoPostel's law: things should parse JSON, but probably only emit SON. That gets you the best of both worlds: you can assume all messages internal to your system are in a canonicalized form (and thus do thing like hashing them with dumb tools), but your system still interoperates with other systems that don't bother to canonicalize, parsing their inputs and sending responses they understand.
- snakeanus 10y ago>making it less human friendly (arguably the thing that made JSON so popular in the first place) I would not call a JSON human friendly due to its lack for comments.
- seagreen 10y agoYou might be interested in JSON5 (http://json5.org/ http://json5.org/) which goes the opposite direction of Son and makes JSON more human-writeable. It's the best, cleanest "JSON+" I've found so far.
- ChuckMcM 10y agoYou might find the discussions around XDR (eXternal Data Representation) helpful in this regard. [1] https://tools.ietf.org/html/rfc4506.html https://tools.ietf.org/html/rfc4506.html
- lobster_johnson 10y agoIs XDR in wide use? Don't think I've ever seen it in the wild. Edit: Apparently used by NFS and ZFS.
- ChuckMcM 10y agoheh, that's funny. Every Linux, FreeBSD, MacOS, and most of the Windows systems. :-) It is how you encode data for ONCRPC. It's fast, it can encode anything, and its dense (no 'convert this to isolatin-8891 first' step). Its also from the 80's so not as visible :-) But perhaps more interesting is the really great machine to machine communication of complex types research that went on in the 80's. Some was interesting (CORBA), some was really scary (ASN.1), and some very fast and minimalist (XDR).
- lobster_johnson 10y agoOh, it's what is also called Sun RPC? I didn't know they renamed it. I guess it's one of the things that fly under one's radar if one isn't actively using a protocol. On Windows a similar niche is (mostly "was") filled by Microsoft RPC, an implementation of DCE RPC, which forms the underlying protocol for DCOM.
- ChuckMcM 10y agoFrom https://www.iana.org/assignments/service-names-port-numbers/service-names-port-numbers.xhtml?&page=3 https://www.iana.org/assignments/service-names-port-numbers/... :-) sunrpc 111 tcp SUN Remote Procedure Call [Chuck_McManis] [Chuck_McManis] sunrpc 111 udp SUN Remote Procedure Call [Chuck_McManis] [Chuck_McManis]
- RodgerTheGreat 10y agoHow about Bencode[1], the data format used in Bittorrent? Values and encodings are already bijective, the format is extremely simple and it's already standardized and described. The only feature "SON" has which Bencode doesn't support is floating-point values, and as you imply above they are very likely to be mutilated in JSON or a subset thereof. [1] https://en.wikipedia.org/wiki/Bencode https://en.wikipedia.org/wiki/Bencode
- seagreen 10y agoThis is the second time Bencode's been mentioned in the context of Son. I didn't know about it before, it definitely looks interesting! However if you're sending data to a service that only accepts JSON then it's obviously not an option.
- RodgerTheGreat 10y agoFair enough. Backwards compatibility is useful in some contexts.
- seagreen 10y agoBy the way, I love that Bencode is bijective. That's a really nice property for a serialization format, and even though I don't mention it on the README it's one of the goals of Son.
- aeijdenberg 10y agoIf you're looking to be able to consistently hash JSON objects you might want to look at Ben Laurie's objecthash: https://github.com/benlaurie/objecthash https://github.com/benlaurie/objecthash It describes a consistent way to hash an object without defining a new format.