6 ms·
Mind elaborating on how the "standard" C++ and Java protobuf implementations are bloated? [edit] I'm genuinely asking. I'm guessing you mean in the generated A
by thisiscorrect 6y ago
Mind elaborating on how the "standard" C++ and Java protobuf implementations are bloated?
[edit] I'm genuinely asking. I'm guessing you mean in the generated APIs, i.e. the code "weight", but maybe you meant something else?
- gabereiser 6y agoProbably a complaint of the codegen and it’s wacky name mangling.
- notretarded 6y agoProbably all this "security" nonsense. What could possibly go wrong with implementing a binary serialisation protocol your self...
- inetknght 6y agoOh look sarcasm! On Hacker News!
- tijsvd 6y agoThings may have improved since, but the implementations are somehow very large and slow. Things may have changed since, but AFAIK the C++ implementation would always allocate on the heap for nested messages, and perhaps even for optional scalars. This may be optimal for larger documents, but not for smallish messages (my use case was market data and trading instructions). I measured certain small messages, where an encode/decode pair would take over a microsecond with Google's implementation, but about 50 ns with a simpler one (versus 15 ns for a memcpy). For Java my experience is mostly with the API itself, which felt very heavy. Edit: I think a lot depends on your use case. I use protobuf mostly as a 'trusted' protocol. If someone didn't set a required field, I don't care. Some bloat may have to do with verifications that I've never needed.
- haberman 6y ago> Things may have changed since, but AFAIK the C++ implementation would always allocate on the heap for nested messages This is no longer the case if you use arenas: https://developers.google.com/protocol-buffers/docs/reference/arenas https://developers.google.com/protocol-buffers/docs/referenc... > and perhaps even for optional scalars This has never been the case, except for string fields where std::string forces us to allocate. Ideally we will eventually use std::string_view for string accessors instead of std::string, so that even string data can be allocated on an arena instead of the heap.
- seppel 6y ago> This has never been the case, except for string fields where std::string forces us to allocate. I'm was quite surprised you didnt offer your own stringview implementation (or something similar) the last time I looked at protobuf. I'd naively assume that inside Google this could be quite a low-efford high-reward optimization.
- haberman 6y ago> I'm was quite surprised you didnt offer your own stringview implementation (or something similar) the last time I looked at protobuf. We sort of do actually: https://github.com/protocolbuffers/protobuf/blob/master/src/google/protobuf/stubs/stringpiece.h https://github.com/protocolbuffers/protobuf/blob/master/src/... The internal version of protobuf lets you switch individual string fields to string_view using [ctype=STRING_PIECE], but migrating the default away from std::string is mainly just an enormous migration challenge. Internally we also do something slightly nuts: we break the encapsulation of std::string so that we can point it to arena-allocated memory (we then "steal" the memory back before the destructor runs). We can only afford to do this internally, where the implementation of std::string is known. The real long-term solution is to move to string_view.
- cbsmith 6y agoIt was never really the case. It's just before arenas it depended on your underlying heap allocator to do all the hard work. Arenas have been around for a long while now though...
- kentonv 6y ago> AFAIK the C++ implementation would always allocate on the heap for nested messages FWIW if you reuse the same message object for multiple parsings, it will re-use the sub-objects as well, thus amortizing away the allocation cost. Parsing the same message into the same object twice should do zero allocations on the second parse. This is the intended way to use Protobuf for small-size messages. Apparently the C++ implementation has also grown support for arena allocation more recently. (After my time, so I don't know much about it.)
- tijsvd 6y agoBut then all that must come with bookkeeping, which brings its own cost. Take a look at an implementation like Prost, for Rust. It's very similar to what I did (10 years ago by now). Everything is just inline, except when messages can be recursive (which should be rare for most protocols).
- kentonv 6y agoThe bookkeeping is not that hard... the pointer is null until it is first allocated, then it remains non-null, while a separate boolean indicates whether the sub-message is actually present in the parent. The big problem is bloat in memory usage if you parse many differently-shaped messages, requiring the app to implement hacks like only reusing a particular object a certain number of times.
- haberman 6y ago> Everything is just inline, except when messages can be recursive (which should be rare for most protocols). Many messages are have lots of optional sub-message fields, and set only a few of them in any given message. These messages would be huge if everything is inline (especially if the same thing happens in those sub-messages). I agree that inlining all sub-messages works great for dense schemas, but it assumes too much about the schema to be a good design for a general-purpose proto library I think. Also maps and repeated fields can never be inline.
- tijsvd 6y ago
- jasonzemos 6y agoSince all other comments appear to contradict you and make apologies from authority (after all, Google can't do anything wrong, right?) I'd like to just reassure you with my 20 years of experience developing C (and the last 10 with C++) in network and system software: that any library -- any library -- that forces internal dynamic memory upon its user smells bad. It's not a universal condemnation, but it begs the question, or rather the skepticism to ask why? In the case of protobufs, there is no real answer. Protobufs is one of the few network libraries that doesn't operate with zero-copy. That steps over a tangible line in the sand. No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications.
- tijsvd 6y ago> No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications. That's a bit harsh. Protobufs deliver smaller wire size than any of the newer "zero-copy" formats. And many receivers of zero-copy formats will... copy the data into some internal representation. If your protobuf implementation delivers classes that are good enough to work with internally (store in maps, forward, etc) then you don't really lose something; instead you gain, due to no manual conversion layer.
- jasonzemos 6y agoCBOR also delivers optimal wire sizes with its variable length encoding. Ceteris paribus, there's no technical advantage to using protobufs. The choice is almost always non-technical (business, platform, partnership, etc) or simple naivete. It's not harsh, it's just engineering.
- haberman 6y ago> make apologies from authority (after all, Google can't do anything wrong, right?) That's not my position at all. In my other comment (https://news.ycombinator.com/item?id=25586447 https://news.ycombinator.com/item?id=25586447) I explain how I've spent 10 years trying to improve on protobuf C++ precisely because I agree that some of these limitations are unnecessary. > No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications. I suppose it depends what you are comparing it to. Almost every JSON library has the same limitations you mentioned, and yet many people find JSON useful for network applications. But I agree that giving users full control over allocations makes a library useful in many more situations. I think arena allocation is a pretty reasonable solution to the problem. You can use whatever memory you want for the arena (stack, heap, static buffer) and you can constrain it so that no heap allocations are allowed. Unfortunately protobuf C++ can't live fully within this arena model while it uses std::string for accessors. Hopefully this can be fixed at some point.