3 ms·
We have already look at CBOR and MessagePack. They miss the ION tables for compact arrays of data. Here is a link to our comparison to other data formats. http:
by VStack 11y ago
We have already look at CBOR and MessagePack. They miss the ION tables for compact arrays of data. Here is a link to our comparison to other data formats. http://tutorials.jenkov.com/iap/ion-vs-other-formats.html http://tutorials.jenkov.com/iap/ion-vs-other-formats.html
- efaref 11y agoI don't really see how ION tables are an improvement over arrays of arrays, e.g.: { "headers": [ "a", "b", "c" ], "rows": [ [ 1, 2, 3 ], [ 4, 5, 6 ], ] } Furthermore, ION appears to require you to know the length of your data up front, whereas you could use CBOR unspecified length arrays to stream data from your database without precalculating the table length. It seems like quite a niche format, though. Most data is not truly tabular. Also, in the table it claims "Yes" under support for "Cyclic references", and yet further down the page: > ION has support for expressing cyclic references between objects. At this point this support is not 100% finalized. So surely this should be "Yes(*)"?
- VStack 11y agoFirst of all, lots of results sent back from backend services (or databases) are arrays of objects. So no, tables are not a "niche" format - tables are heavily used. Second, an array of arrays could mean anything. You have not semantics telling whether the arrays are independent or if the first array is an array of columns for the following arrays. ION Tables add that semantic information. Third, yes, we could encode everything as text or as raw bytes and leave it up to the user to make sense of it. But that is exactly what we are trying to avoid with ION. We want to give devs a decent standard data format to use, that doesn't require a lot of data encoding choices up front. The encoding options have been thought through already, and sensible choices already made which you can just follow. Fourth, you can nest ION tables inside ION tables, and thus create a more compact representation of an object graph. Using JSON / CBOR you would need a lot of nested arrays inside arrays to emulate that. Possible, but not exactly pretty.
- VStack 11y agoYes also, the TLV nature of ION means that you have to buffer the ION data while writing it, in case you don't know the full size of the embedded ahead of time. However, this is really only a problem for very big messages. HTTP has worked like that for a long time already, and HTTP servers have proven capable of scaling pretty well, won't you agree? Additionally, if a message does not include its length up front, you are just trading faster write time for slower read time. A node receiving dynamically sized data then has the problem of knowing how much data to allocate for the full message, plus the receiver has to inspect the data as it comes in to see where it ends. With an ION message the receiver knows within the first 4-5 bytes (typically) how big that message will be, and can thus copy the following bytes directly into the perfectly allocated memory area, without having to examine them any further. Since data is most often read more times than written (e.g. data written to a file), we felt that making a tradeoff that favors read speed over write speed made sense.
- brianolson 11y agoIf you're going to send structured records and care about knowing the structure: use protobuf. If you're going to be binary JSON: use RFC 7049 CBOR. http://cbor.io/ http://cbor.io/
- room271 11y agoHave you also looked at Transit? https://github.com/cognitect/transit-format https://github.com/cognitect/transit-format
- VStack 11y agoYes, we have looked at Transit recently. Transit encodes using JSON or MessagePack. We believe ION to be a more versatile data format than MessagePack, so Transit could benefit from serializing to ION instead of MessagePack.
- krschultz 11y agoI don't see Apache Avro in your comparison table. Have you looked at that?
- VStack 11y agoYes, we have looked at both Apache Avro and Thrift. The ION encoding is similar to these. Avro uses schemas, just like Protobuf, so it is not so easy to route for intermediate nodes. Thrift uses its own IDL schema language too.