4 ms·
I played quite a bit with MessagePack, used it for various things, and I don't like it. My primary gripes are: + The Object and Array needs to be entirely and
by buserror 2y ago
I played quite a bit with MessagePack, used it for various things, and I don't like it. My primary gripes are:
+ The Object and Array needs to be entirely and deep parsed. You cannot skip them.
+ Object and Array cannot be streamed when writing. They require a 'count' at the beginning, and since the 'count' size can vary in number of bytes, you can't even "walk back" and update it. It would have been MUCH, MUCH better to have a "begin" and "end" tag --- err pretty much like JSON has, really.
You can alleviate the problems by using extensions, store a byte count to skip etc etc but really, if you start there, might as well use another format altogether.
Also, from my tests, it is not particularly more compact, unless again you spend some time and add a hash table for keys and embed that -- but then again, at that point where it becomes valuable, might as well gzip the JSON!
So in the end it is a lot better in my experience to use some sort of 'extended' JSON format, with the idiocies removed (trailing commas, forcing double-quote for keys etc).
- masklinn 2y ago> Object and Array cannot be streamed when writing. They require a 'count' at the beginning Most languages know exactly how many elements a collection has (to say nothing of the number of members in a struct).
- touisteur 2y agoI think the pattern in question might be (for example) the way some people (like me) sometimes write JSON as a trace of execution, sometimes directly to stdout (so, no going back in the stream). You're not serializing a structure but directly writing it as you go. So you don't know in advance how many objects you'll have in an array.
- deleted 2y ago[deleted]
- buserror 2y agoIf your code is compartimented properly, a lower layers (sub objects) doesn't have to have to do all kind of preparations just because a higher layer has "special needs". For example, pseudo code in a sub-function: if (that) write_field('that'); if (these) write_field('these'); With messagepack you have to go and apply the logic to count, then again to write. And keep a state for each levels etc.
- masklinn 2y agoThat sounds less than compartmentalisation and more like you having special needs and being unhappy they are not catered to.
- atombender 2y agoNot if you're streaming input data where you cannot know the size ahead of time, and you want to pipeline the processing so that output is written in lockstep with the input. It might not be the entire dataset that's streamed. For example, consider serializing something like [fetch(url1), join(fetch(url2), fetch(url3))]. The outer count is knowable, but the inner isn't. Even if the size of fetch(url2) and fetch(url3) are known, evaluating a join function may produce an unknown number of matches in its (streaming) output. JSON, Protobuf, etc. can be very efficiently streamed, but it sounds like MessagePack is not designed for this. So processing the above would require pre-rendering the data in memory and then serializing it, which may require too much memory.
- baobun 2y ago> JSON, Protobuf, etc. can be very efficiently streamed Protobuf yes, JSON no: you can't properly deserialize a JSON collection until it is fully consumed. The same issue you're highlighting for serializing MessagePack occurs when deserializing JSON. I think MessagePack is very much written with streaming in mind. It makes sense to trade write-efficiency for read-efficiency. Especially as the entity primarily affected by the tradeoff is the one making the cut, in case of msgpack. It all depends on your workloads but Ive done benchmarks for past work where msgpack came up on top. It can often be a good fit for when you need to do stuff in Redis. (If anyone thinks to counter with JSONL, well, there's no reason you can't do the same with msgpack).
- atombender 2y agoSorry, I was mentally thinking of writing mostly. With JSON the main problem is, as you say, read efficiency.
- laurencerowe 2y agoThe advantage of JSON for streaming is on serialization. A server can begin streaming the response to the client before the length of the data is known. JSON Lines is particularly helpful for JavaScript clients where streaming JSON parsers tend to be much slower than JSON.parse.
- baobun 2y agoThe topic is data serialiasation formats, not programming languages.
- crabmusket 2y ago> trailing commas, forcing double-quote for keys etc How do these things matter in any use case where a binary protocol might be a viable alternative? These specific issues are problems for human-readability and -writability, right? But if msgpack was a viable technology for a particular use case, those concerns must already not exist.
- nullfield 2y agoI think this is the point-when one wanted “easy” parsing and readability for humans, they abandoned binary protocols for JSON; now, people who are finding some performance issue they don’t like are trying to start over and re-learn all the past lessons of why and how binary protocols were used in the first place. The cost will be high, just like the cost for having to relearn CS basics for non-trivial JS use was/is.
- your_fin 2y agoI must further prothelitize Amazon Ion, which solves for most of the listed complaints and is criminally underused: https://amazon-ion.github.io/ion-docs/ https://amazon-ion.github.io/ion-docs/ The "object and array need to be entirely and deep parsed" and "object and array cannot be streamed when writing" are somewhat incompatible from a (theoretical) parsing perspective, though; you need to know how far to skip ahead in order to do so. I agree that it is silly to design an efficiency-oriented format that does neither, though. Ion chooses to be shallow parsed efficiently, although it also makes affordances for streams of top-level values explicitly in the spec.