3 ms·
I have been thought all the serialization formats such as Protobuf, Thrift, BSON, or MessagePack are using compression as much as possible because saving amount
by eonil 13y ago
I have been thought all the serialization formats such as Protobuf, Thrift, BSON, or MessagePack are using compression as much as possible because saving amount of I/O is ultimate win for overall performance rather than fast calculation by memory alignment.
When I see this Capnproto, I am confusing that which one is right approach. Alignment is not exotic technique, and why didn't they align the data if there's no reason to save I/O?
Or am I totally misunderstanding these implementations?
- kentonv 13y agoWell, it depends on the environment. If you are doing interprocess communication, then I/O bandwidth is obviously not a concern at all. On the other hand, over the internet, it clearly is the biggest concern. For intra-datacenter traffic on a 10Gbit NIC, it's harder to say, but it _probably_ isn't the bottleneck. Cap'n Proto supports both cases by making additional packing optional, so you can choose the best trade-off for your application. Regarding the other formats you mention, I think you may be imagining that the designers of these protocols thought more carefully about them than they really did. Protobuf, for example, was designed pretty ad-hoc to solve an immediate problem in Google's search infrastructure, and then stuck mostly because as more and more things used it, it was easier to keep using it than start over. The designers readily acknowledge that it is not an ideal format -- in fact, there are other ways they could have done the encoding which would have taken no more space but would have saved significant CPU time. (Disclosure: I was the maintainer of protobufs for a long time, though not the original creator. I am also the author of Cap'n Proto.)