5 ms·
Dealing with data as my day job I've become highly sensitive to schema's. And I hate those "schemaless" (aka schema-on-read) serialization formats more and more
by rollulus 8y ago
Dealing with data as my day job I've become highly sensitive to schema's. And I hate those "schemaless" (aka schema-on-read) serialization formats more and more. No, there is no schemaless, there is a schema, but it is buried in your code in a convoluted way on each line where you read and interpret your deserialised data and all tests and assumptions you have there are a horrible representation of your schema. That is an important dimension to compare serialization on, and JSON is a complete loser in that sense. Not to mention how it flawed its numeric type is.
- josephg 8y agoI really wish we had a JSON 2 format which could fix JSON's obvious shortcomings. I would like to see: - Support for Maps and Sets (unlike objects, maps allow arbitrary types to be used as keys) - A standard Date format - Embedded binary blobs. No idea how to do this and keep it human readable, but when you need this its super useful. Maybe something similar to WS's binary message encoding. - Arbitrary precision integer support. This is particularly useful for cryptocurrencies and for interoperability with 64 bit integers in other languages. And bigints are coming to javascript - https://github.com/tc39/proposal-bigint https://github.com/tc39/proposal-bigint - Maybe even fix JSON's weird unicode encoding: http://timelessrepo.com/json-isnt-a-javascript-subset http://timelessrepo.com/json-isnt-a-javascript-subset Unfortunately it seems like nobody 'owns' JSON enough to give JSON 2.0 the political weight it would need for cross-language support.
- dchest 8y agoArbitrary precision integers is a JavaScript implementation detail; JSON standard doesn't specify number precision.
- josephg 8y agoHuh good to know: $ node > JSON.parse('1231231231231231231231231123123123123') 1.2312312312312312e+36 I've heard of JSON implementations in other languages making bigger JSON numbers than javascript supports; but I didn't realise JSON.parse would quietly parse them and throw away precision in the process. I wonder what the plan is for JSON serialization support of bigints. I assume there is no plan - Chrome stable 67 supports bigints, but JSON.serialize(2n) throws a TypeError.
- sharkdesk 8y agoIn theory, that is true. In practice, if the consuming client is javascript, it is really hard to map the numbers into an arbitrary precision number upon decoding (which would come from some library like big.js), rather than a pure double. Which means, in practice, you are stuck passing numbers as strings (and on the javascript side, pass that string into the arbitrary precision number library)
- pitaj 8y agoMaps and Sets are entirely irrelevant when it comes to serialization. You may as well store or transmit that information as an array (of pairs, in case of Map). The point of a Map or Set (O(1) insertion / removal / contains) don't matter when you're talking about serialization.
- repsilat 8y agoWell, unlees you're using something like capnp, where O(1) contains is meaningful when O(n) deserialise is not acceptable. And really, if you have consumers/producers in different languages/codebases, built-in support for these things is a great convenience. You wouldn't say that objects/dicts in JSON are irrelevant, would you?
- derefr 8y ago> You wouldn't say that objects/dicts in JSON are irrelevant, would you? I kind of would. Coming from Erlang, I don't see a point to there being a whole separate syntax for encoding object/map types. Just encode objects/maps as arrays of pairs (2-Tuples, or just length-2 arrays). Then, if you want ser-des type fidelity, stick an annotation onto the array (like in YAML) to say what type it should decode to. E.g., something like: [1, 2, 3] # array [['a', 1], ['b', 2], ['c', 3]] # array of pairs [@object, ['a', 1], ['b', 2], ['c', 3]] # array of pairs hinted so it should decode to an Object [@map, ['a', 1], ['b', 2], ['c', 3]] # array of pairs hinted so it should decode to a Map The nice thing about this approach is that JSON libraries could do as much or as little work in parsing out the annotations as they want: they could recognize the annotations and construct the referred-to type; or they could just pass back the array with the annotation expressed as an Annotation value.
- tonyg 8y agoDo you want remote code execution? Because that's how you get remote code execution. http://docs.couchdb.org/en/2.1.1/cve/2017-12635.html http://docs.couchdb.org/en/2.1.1/cve/2017-12635.html Map and set values are important for many applications. The fact that JSON doesn't have a way of denoting a map or set value (or anything else, but that's another issue) is a problem: it means there's no understanding common to all JSON consumers about what syntax denotes a map or a set. Shortcuts are taken, consumers disagree on the details, and the linked CVE is the result. Being able to reliably convey "this is intended to be a map", or "this is intended to be a set" is crucial for secure and robust interoperability.
- tlocke 8y agoYou might be interested in Zish https://github.com/tlocke/zish https://github.com/tlocke/zish which has support for maps, timestamps, binary data and decimal types.
- teekert 8y agoI'm not an expert, just a bio-data-scientist learning every day... But wouldn't Yaml be what you are looking for?
- pletnes 8y agoJson is for machines, yaml is for humans, as a general rule. They’re mostly compatible feature-wise.
- skolemtotem 8y agoAnd if I remember correctly, YAML is a superset of JSON.
- teekert 8y agoYaml is well readable by a human but it's much to tightly defined to be "for humans". It's for machines but readable by humans I'd say. Plus it's got sets and you can embed csv's (just 2 things of the top of my head).
- tr33house 8y agoyou want to use protobufs
- strkek 8y agoObligatory link to the "YAML sucks" repo: https://github.com/cblp/yaml-sucks https://github.com/cblp/yaml-sucks If you ignore the flame-inducing title, it's just a table showing how different implementations parse YAML input in very different ways.
- phyrex 8y agoDo you know of Transit? https://github.com/cognitect/transit-format https://github.com/cognitect/transit-format
- randyrand 8y agoNo need to extend json itself. But the libraries would need to be extended. maps can be implemented with 1 array and 1 map, with the keys being the hash of the object. the hash function should probably be written in the Json itself for completeness. embedded beinary blobs already work. lookup GLB for an example on how to embed binary blobs in Json. as others said, json already supports arbitrary precision.
- derefr 8y ago> embedded beinary blobs already work. lookup GLB for an example on how to embed binary blobs in Json. Those aren't embedded binary blobs, those are string representations of base64-encoded binary blobs. Unless they come out of the JSON decoder as a byte array, they're not "working" as part of JSON; they're another standard on top of JSON. (Also, the GP commenter probably wants them to be transmitted with 0 encoding overhead. Which can't really work while JSON is still JSON. But you can always use an alternative format which handles a superset of JSON's types, and which is binary, e.g. BSON or CBOR.) > maps can be implemented with 1 array and 1 map, with the keys being the hash of the object. the hash function should probably be written in the Json itself for completeness. I think you're fundamentally misunderstanding the thrust of the GP poster's request, here? They don't want to serialize a map in a way that is cheap to deserialize—different languages have different hashtable implementations so there's no way it could really work. What the GP poster (and many other people) want, is just to have something that encodes similarly to existing "object maps" (the ones with curly braces), but with a slight syntactic alteration so that they come out of the decoder as Maps, rather than as Objects.
- randyrand 8y agoI'm not talking about data URIs. You can write binary data in the same file as the json. Most parsers don't care what comes after the last curly brace.
- aepiepaey 8y agoAnother couple of missings: - NaN/inf floats - Comments
- tr33house 8y agoprotobufs already support all these
- adamdrake 8y ago1000x yes! Schemas are defined somewhere, even if it's a poor and bug-ridden definition in the code. Ditto for the point that JSON doesn't count. I wrote an article, On Schemas, to try and elaborate on this topic a bit. https://adamdrake.com/on-schemas.html https://adamdrake.com/on-schemas.html
- scarface74 8y agoI'm assuming you are using a weakly typed language. My schema is defined by my data objects and I use the built in validation attributes in C#. If the request doesn't pass the validation part of the pipeline, my controller never gets the request and a Bad Request message is sent back to the client.
- Cthulhu_ 8y agoI keep reiterating that there was nothing wrong with XML.
- teddyh 8y agoOr ASN.1.
- zeveb 8y agoOr S-expressions.
- tonyg 8y agoBoth ASN.1 and XML are standardized (with extensive, lengthy and detailed specifications!). The term "S-expression" covers a very broad family of vaguely related not-very-interoperable surface syntaxes.
- lisper 8y agoS-expressions are standardized in both Common Lisp and Scheme. The two standards are not identical, but are mostly compatible, with the CL standard being mostly a superset of the Scheme standard.
- tonyg 8y agoYes, that's exactly what I mean. It's like saying "we use JSON as our file format", but worse. > "So this program saves its data as 'S-expressions'." > "Oh, cool. Which dialect?" > "Uh, I don't know. It doesn't say in the README. It just says 'S-expressions'." The answer can be CL, Scheme (lots of variations and dialects even here), the never-finished SPKI Sexps, OCaml sexps, something that the new dev on the team cooked up last Tuesday that vaguely resembles what they learned as "lisp" in college, or something else entirely. On the other hand, if one were to say "this program uses R4RS S-expressions" (or, presumably, "CL S-expressions", but I haven't read the relevant bits of CLtL), you'd immediately be in a much nicer place than JSON can offer. Not only would you have a well-specified syntax for a reasonably broad range of data types, you'd have a useful equational theory as well. [ETA: Unless you want unicode. Doh. R6RS, maybe.] Ah, the impossible dream.
- makmanalp 8y ago> I hate those "schemaless" (aka schema-on-read) serialization formats more and more This, and also I think a lot of people who preach the flexibility of schemaless anything really just wants a more flexible schema definition language, with a lot of "maybe" options and similar stuff.
- deleted 8y ago[deleted]