8 ms·
> I highly encourage any greenfield project to look into well designed and better specified alternatives. Like what?
by kortex 5y ago
> I highly encourage any greenfield project to look into well designed and better specified alternatives.
Like what?
- jerf 5y agoPart of the problem is that there's at least half-a-dozen high quality answers out of the gate (gRPC, FlatBuffers, Protocol Buffers, XML in some cases, Thrift), and an even-longer long tail after that. It's made harder when four different teams who deeply loathe JSON and independently decide to use something "better" can legitimately use four completely different technologies if they don't communicate with each other.
- 35fbe7d3d5b9 5y agoTo your comment above – you can bodge around interop problems with JSON in ways that you cannot with some of these other technologies. I like to joke that I invented ndjson over a decade ago when I accidentally forgot to put things in an array before `json.dumps`, I just wasn't smart enough to call it a standard. But when you do end up with ndjson when you wanted an array of results, or vice versa, JSON makes it easy to munge things to where you need. Compare that to something like protobuf: it's not a self-synchronizing stream, so if you send someone multiple messages without framing them (prefix by length or delimited are popular approaches), they're going to decode a single message that doesn't make much sense on the other end. And they won't be able to fix it at all. So I guess JSON is New Jersey style design[1]. [1]: https://dreamsongs.com/RiseOfWorseIsBetter.html https://dreamsongs.com/RiseOfWorseIsBetter.html
- kortex 5y agoWell, you invented one of the best things since sliced bread! I love NDjson, being able to parse a sequence of {} objects as an array is just frankly more natural. A coworker got some absurd speedup going from some massive json array to ndjson. Honestly if json had as part of its spec line-delimited arrays, and accepting NaN, it'd be close to perfect. Oh and native ints, but that is JS's problem. Well, and a single, canonical spec. And a hard limit (however high) on nesting depth. And some other things. Ok, maybe it's far from perfect.
- q3k 5y ago> Compare that to something like protobuf: it's not a self-synchronizing stream, so if you send someone multiple messages without framing them (prefix by length or delimited are popular approaches), they're going to decode a single message that doesn't make much sense on the other end. And they won't be able to fix it at all. FWIW, this is a conscious design decision with Protobuf: it allows for easy upsert operations on serialized messages by appending another message with the updated field values. This is very useful for middleware that wants to either just add its own context to a message it doesn't even parse [1], or for middleware that might handle protobuf messages serialized with unknown fields. On the other hand, 'newline delimited protobuf' is much less useful day-to-day than ndjson, as gRPC provides message streaming, which solves the issue of wanting to stream small elements of a long response (which is the general usecase of ndjson from my experience). For on-disk storage of sequential protobufs (or any other data, really), you should be using something like riegeli [2], as it provides critical features like seek offsets, compression and corruption resiliency. [1] - eg. passing a Request message from some web server frontend, through request routers, logging, ACL and ratelimit systems up to the actual service handling the request. [2] - https://github.com/google/riegeli https://github.com/google/riegeli
- syncsynchalt 5y ago> teams who deeply loathe JSON In the current world this seems like a lifestyle choice that sets yourself up for constant self-punishment. I might be a curmudgeon but I'll take JSON for data interop any day over anything that _requires_ tooling (protobuf, gRPC). And I'll take it over the XML ecosystem too. The faults of JSON seem, in practice, to be less harmful than the faults of other formats.
- q3k 5y agoMy preference is Protobuf, but really anything that's not JSON and which also comes with some IDL gets my approval.
- kortex 5y agoI like protobuf for some use-casess (namely grpc) but a) it's a binary format and sometimes (often times) it's nice to have a text protocol b) protobuf libraries and protoc have given me way more grief overall than json (python, js, c++) If your workflow already supports it, I can see it being useful, but it's got a pretty steep learning curve to be honest, certainly more than json, despite the ill-implemented libs out there. If I wanted a binary format, IMHO I'd go for msgpack first, and reach for protobuf if that didn't work for me.
- elteto 5y ago> I like protobuf for some use-casess (namely grpc) but a) it's a binary format and sometimes (often times) it's nice to have a text protocol Protobuf (and flatbuffers) supports parsing messages from JSON instead of a binary blob. Best of both worlds IMO.
- avmich 5y agoCan you use JSON Schema? Generating classes from it, if you want native objects?
- q3k 5y agoYou can use whatever you want :). I personally would rather still go with Protobuf if I'm going to put in the effort to add a schema and codegen. It gives me other nice-to-have features (faster [de]serialization, smaller messages, field numbers and schema evolution, nicer IDL [not JSON!], gRPC, ...) and does away with some problems intrinsic to JSON that no schema system will fix (terrible number type, lack of binary type, slow parsing). It also has some interop with JSON in the rare case you absolutely positively need to convert to/from it (which is IMO the only upside of using JSON Schema in case you need that interop).
- 0xbadcafebee 5y agoYAML. Of course implementations of this go all over the place too, but you could say the same of XML parsers to a certain extent. I still pine for binary-only data formats. They're easier to program, and nobody makes the mistake of trying to edit them manually or compose them in a shell script. Parsing data shouldn't be hard, but it also shouldn't be so easy that people hang themselves by accident. Of course, the reason why we largely have text data formats is because it's insanely simpler to troubleshoot systems that use them. Some things should just be easier to manipulate. But for general purpose work, I miss binary data formats. Zip is probably my favorite general-purpose binary data format. It's old, well defined, works with any kind of data, and you can immediately seek to data in very large archives rather than having to parse the entire thing first. And then there's that whole compression thing. If you wanted to distribute a thousand tiny blobs of CSV, JSON, YAML, and XML, all in one container, you could do much worse than Zip.
- rjh29 5y agoI've had a number of negative experiences with yaml, enough to put me off using it. For example the implicit parsing of 'yes' and 'no' into bools rather than strings (including the NO country code for Norway) <https://hitchdev.com/strictyaml/why/implicit-typing-removed/ https://hitchdev.com/strictyaml/why/implicit-typing-removed/>, the no-quote rules allowing accidental creation of inline hashes/arrays <https://hitchdev.com/strictyaml/why/flow-style-removed/ https://hitchdev.com/strictyaml/why/flow-style-removed/>, multiline string syntax so complex that it needs a helper tool <http://yaml-multiline.info/ http://yaml-multiline.info/>, and powerful extensions that invite your program to be exploited <https://www.sitepoint.com/anatomy-of-an-exploit-an-in-depth-look-at-the-rails-yaml-vulnerability/ https://www.sitepoint.com/anatomy-of-an-exploit-an-in-depth-...> It manages to be both a poor data interchange language compared to JSON, and also a bad human-friendly langage due to the above ambiguities. Unfortunately it's still the best human-friendly configuration language in wide use, so I use strictyaml (https://hitchdev.com/strictyaml/ https://hitchdev.com/strictyaml/) instead.
- 0xbadcafebee 5y agoSorry, I should have been more explicit: YAML should never, ever, be edited by a human. Neither should JSON. But they can be read by a human, and copy+pasted, which makes them easier to troubleshoot. All of the problems you list of YAML are due to humans making assumptions, rather than having a program do all of the serialization/deserialization. YAML has a ton of useful data structures and they do not cause problems when you use a real parser to populate or interpret them. It's also a superset of JSON, so it's clearly not poor in comparison to JSON. It's also not a configuration language. It's a data serialization format.
- NavinF 5y agohttps://capnproto.org/ https://capnproto.org/