5 ms·
For context, let's start with what JSON is. All the JSON spec really lets you do is say "this sequence of code points is JSON", or "this sequence isn't". It inc
by seagreen 9y ago
For context, let's start with what JSON is. All the JSON spec really lets you do is say "this sequence of code points is JSON", or "this sequence isn't". It includes a few interoperability suggestions for how you might avoid JSON that will be hard to interpret, but it's extremely stark on parsing and generation.
For a pretty epic summary of this, see this Reddit comment: https://www.reddit.com/r/programming/comments/59htn7/parsing_json_is_a_minefield/d98qxtj/ https://www.reddit.com/r/programming/comments/59htn7/parsing...
You have a comment downthread about how inadequate this is, which I totally agree with:
> The behaviour of the encoder and decoder are just as important to specify as the bytes across the wire. This is the whole problem with JSON in the first place.
Unfortunately since the goal of Son is to be a subset of JSON, there's not much we can do about encoding/decoding within Son's spec. Instead I hope to build a sane foundation on which other tools can be built. Limiting the scope of Son seems like a good way to go about this. In this case, that means not specifying ranges for numbers, so that future tools such as schemas have as much flexibility as possible (for instance to include arbitrary precision numbers).
EDIT: I'm mixing together two different things in my last paragraph. I don't want Son to know about Floats, Doubles, etc since JSON doesn't. That doesn't mean that we couldn't specify a range on number size though. The reason that is left out is to allow maximum flexibility for other tools to build on Son -- I'm trying to drop extraneous parts of JSON, but as few useful parts as possible. This needs more elaboration within the project's docs.
- ghettoimp 9y ago"In this case, that means not specifying ranges for numbers, so that future tools such as schemas have as much flexibility as possible (for instance to include arbitrary precision numbers)." If the allowed range of numbers isn't specified, then when you emit data in this format, you can't be sure it will be read correctly on the other end...
- seagreen 9y agoThis is definitely a problem, but the place to solve it isn't Son. Son is intended to be a very simple project that grabs all the clear wins. Do we really need both `0` and `-0`, stuff like that. There are too many different things JSON Numbers can represent for us to have a clear strategy for all them in Son. Int64s, doubles, floats, arbitrary precision numbers (eg https://hackage.haskell.org/package/scientific-0.3.5.2 https://hackage.haskell.org/package/scientific-0.3.5.2), etc. That makes this a good for for the schema layer (eg JSON Schema or whatever you're using), not the base specification layer.
- patrec 9y ago> Do we really need both `0` and `-0`, stuff like that. Well, yeah since signed zeros are a thing. Chrome: > JSON.parse('-0.0') -0 > JSON.parse('0.0') 0 python: In [3]: json.loads('-0.0') Out[3]: -0.0 In [4]: json.loads('0.0') Out[4]: 0.0
- patrec 9y ago> This is definitely a problem, but the place to solve it isn't Son. But then what problem does Son actually solve? How is a canonicalisation format that's not actually canonical useful? You take a hit (e.g. by forgoing fast existing libraries, control over pretty printing, and by being forced to sort on serialization which screws up streaming), but gain essentially nothing, not even the advertised benefit of your stuff not randomly changing as it moves through the stack. The likely main benefit of canonicalization is for security and crypto, but unless you use son without numbers (or create your own subset of son) it's kinda useless for that. Of course you can just encode all your numbers in strings or something, but at the point where you're doing your own number parsing logic (which is the hardest bit), why not just use some well-designed actually canonical format like csexp?