10 ms·
Protobuffers Are Wrong (2018)
- rektide 4y agoThe type system woes are quite legit. I wonder how many of these were improved in FlatBuffers. Which has a ton of other good things going for it, chiefly ideally less serialization demands. Once you start doing rpc, the demand goes so up. Cap'n'Proto's ability to return & process future values is such an exciting twist, that melds with where we are with local procedure calls: they return promise objects that we can pass all around while work goes on. Kenton popped up a couple months ago to mention maybe finding time for mutli-party support, where I can for ex request say the address of a building & send the future result to another 3rd party system. Now this is less just a way to serialize some stuff, & more a way to imagine connecting interesting systems.
- the-alchemist 4y agoIs Cap'n'Proto [0] as awesome as it sounds? Anyone have any first-hand experience and can speak on it? Thanks! [0]: https://capnproto.org/ https://capnproto.org/
- ChadNauseam 4y agoI've used Cap'n'Proto for a work project, and aside from not-always-amazing language compatibility it really is awesome. It's what I wish Protobufs was. The performance improvements don't matter much to me, the biggest thing for me is the type system fixes most of the mistakes that are mentioned in this article.
- avinassh 4y agois it possible to use this with gRPC? edit: seems, not possible as of now - https://github.com/capnproto/capnproto/issues/478 https://github.com/capnproto/capnproto/issues/478
- avinassh 4y agoI was looking to Cap'n Proto's page and they say: > Cap’n Proto gets a perfect score because there is no encoding/decoding step. How do they achieve this?
- zarzavat 4y agoIt’s an untruth. There is a validation step of course, which means that there is a decoding step. There is no encoding step because the wire format is the same as the in-memory representation.
- kentonv 4y ago> There is a validation step of course No, there isn't. Not any more than there is a decode step anyway. The truth is that encoding, decoding, and validation all occur when you call the accessor methods for specific fields. If you never actually call the getter for some field, it won't be decoded nor validated at all. If you are going to be reading the entire message tree, then zero-copy is only an incremental improvement on something like Protobuf -- the use of fixed-width values may make decoding faster (or slower, in some cases). The real magic is when you want to read just one field of a 10GB file. Just mmap() the whole thing and treat it like a byte array, and it'll be efficient. Cap'n Proto will not attempt to scan the whole file, only the parts that you explicitly query. (I'm the author of Cap'n Proto.)
- jsnell 4y agoSignificant previous discussions: https://news.ycombinator.com/item?id=21871514 https://news.ycombinator.com/item?id=21871514 (211 comments) https://news.ycombinator.com/item?id=18188519 https://news.ycombinator.com/item?id=18188519 (298 comments) In particular, Kenton's rebuttal at https://news.ycombinator.com/item?id=18190005 https://news.ycombinator.com/item?id=18190005 is worth a read.
- taeric 4y agoI was amused to find I had, in fact, already discussed this one. :D
- morelisp 4y agoprotobufs are a data exchange format. The schema needs to map clearly to the wire format because it is a notation for the wire format, not your object model du jour. If the protobuf schema code generator had to translate these suggestions into efficient wire representations and efficient language representations, it would be more complex than half the compilers of the languages it targets. I do mourn the days when a project could bang out a great purpose-built binary serialization format and actually use it. But half the people I hire today, and everyone in the team down the hallway that needs to use our API, can no longer do that. I'm lucky if they know how two's complement works
- mytailorisrich 4y agoWe use C and all our nodes are x86... so we just dump packed structs on the wire and the schema is a header file. I suspect this sort of simple approach works in most cases... And really that's how the network stack on Linux and Co. is implemented (modulo taking care of endianness)
- thadt 4y agoYep, that does work quite well in a fixed scope. Till one day someone wakes up and wants a UI in C#. Which is also not a big deal, but you have to either hand roll the serialization code or build a header file parser and generate it. And then someone needs to add a field - so we just tack it onto the end of the struct. This works fine as long as you have the length represented out of band somehow. If not then you get to make struct FooBar2 with your extra field. Then a year later someone has the great idea that it would be great to send text, and now you're making a variable length packet or just always sending around structs with a bunch of empty space and a length field (which is also not too bad). But wait, now the UI team totally needs that data in JavaScript for their new Web UI, so you're back to either generating or hand jamming dozens of structs. All the while tamping out bugs in various places where someone is running old code against new data or old data against new code - all of which have various expectations. Or that microcontroller that is blowing up because it just doesn't like that unaligned integer that is packed into the struct. Not that I'm bitter about that life - it just involved a lot more troubleshooting protocols over the years than I'd have liked. Anyway, those are problems that Protocol Buffers helps solve. But as long as you're using just C and not changing much - packed structs are quite lovely.
- paxys 4y ago> They’re clearly written by amateurs Protocol buffers were written by Jeff Dean and Sanjay Ghemawat. Whether you see issues with the implementation or not and whether they are fit for your use case or not are valid discussions, but if your argument starts with "omg these idiots couldn't design a simple product" then I'm already reading the rest of it with a huge grain of salt.
- grumpyprole 4y agoI think the emotional response is probably due to the fact that Protocol buffers is often forced upon teams, due to the desire for a common standard and the marketing power that Google have. But the author raises many valid issues and you would be wrong to dismiss them completely. The first version of Protocol buffers, IIRC, didn't even support static offsets and extracting a field without deserialising the entire message. So it was arguably both inelegant and inefficient. I was never impressed myself personally and have seen quite a few better closed-source wire formats inside various corporates.
- paxys 4y agoYes you can have many valid issues with Protocol buffers. "Written by amateurs" is not one of them, and that is the specific one I was dismissing. Reading through the rest of the post though, it is pretty clear that it is a case of trying to use what is a standard for data transfer over the network to describe your entire application's object model. It isn't the right tool for the job, and doesn't have to be. If it is being forced on you then, well, that's a complaint for your management, not the technology itself.
- bragr 4y ago>I think the emotional response is probably due to the fact that Protocol buffers is often forced upon teams True, but a lack of self awareness is never a good look.
- hot_gril 4y agoIt's also a really bold claim that he doesn't back up, so I don't trust the rest. It's not like he dug around and found out that the main protobuf author was actually some hack who forced something bad on the rest of the company then bailed. That would be Google Plus.
- shaftway 4y agoPersonally I avoid Map fields in protos. I don't find a huge amount of value for a Map<int, Foo> over a repeated Foo with an int key. That tends to be more flexible over time (since you can use different key fields) and it sidesteps all of the issues about composability of them. I think required fields are fine, provided that you understand that "required" means "required forever". If you're already using protos this isn't exactly a brand new concept. When you use a field number that field number is assigned forever, you can't reuse it once that field is deprecated. It requires a bit more thoughtful design, and obviously not all fields should be required, but it has value in some places.
- summerlight 4y agoThere are some legit criticisms here (remember, protobuf is a more than 20 years old format with extremely high expectation of backward compatibility so there are lots of unfixable issues), but in general this article reveals a fundamental misunderstanding of protobuf's goal. It is not a tool for elegant description of schema/data, but evolving distributed systems composed of thousands of services and storage. We're not talking about just code, but about petabytes of data potentially communicated through incomprehensible level of complex topology.
- rektide 4y agoSemi-subtle pitch for worse-is-better, the constraints & limitations are what make it work in this chaotic difficult gnarly uber-legacy micro-verse.
- hot_gril 4y agoYeah, I stopped reading at "by far, the biggest problem with protobuffers is their terrible type-system" for similar reasons. What did the author want, ASN.1? Cause that gives you a full type system, except that's not what I want. Edit: kentonv's rebuttal says this too. "This article appears to be written by a programming language design theorist who, unfortunately, does not understand (or, perhaps, does not value) practical software engineering." And when my angry hat is on, this is also how I would generalize programming language design theorists.
- mianos 4y agoFunny you should mention ASN.1. Protobuffers is Really just a toy ASN.1 written by some guys who just solved the problem at hand instead of looking further afar. It's great for what it is. Many of us have written something similar. It's not a pro level system designed by a serious telecom team.
- hot_gril 4y agoI worked on LTE stuff during a research/engineering position. It was just 1 year and not my life's work, but I had to deeply understand ASN.1 for it, and I'd never heard of Google's Protobuf. I didn't like ASN.1 even for its intended use case. That big type system was far more complex than necessary for the protocols I was dealing with (surrounding the EPC mainly). ASN.1 is more optimized in other ways that make sense, sacrificing some flexibility in favor of performance because the ends are expected to be more coordinated, but that's separate.
- theamk 4y agoThis was hard for read because of the glaring omissions of existing engineering practices. It feels author only cares about programming purity and writing programs which transform metadata. Making all fields in a message required makes messages into product types but loses compatibility with older or newer versions of protocol. Auto-creating objects on read dramatically increases chances that a field removal in future version of protocol will be handled properly. Auto-creating on write simplifies writing complex setters, etc... You could argue that those problems could be solved better (and I might agree, protobufs are one of my least favorite serialization protocols) but not even acknowledging reasons and pretending the designers didn't know better makes writer seem either ignorant or arguing in bad faith.
- bluGill 4y agoI find rolling my own serialization protocol easy enough, and I want some friction from using it: i've been burned when it is too easy and people start to use IPC where a function call would just as well and be more efficient. Premature pessimization is a large problem these days.
- jchw 4y agoProtobuf is the worst serialization format and schema IDL, except for all the others. The truth is, a lot of these critiques are valid. You could redesign protobuf from the ground up and improve it greatly. However, I think this article is bad. Firstly, it repeatedly implies that protobuf was designed by amateurs because of its design pitfalls, ignoring the fact that it's also clear that protobuf grew organically into what it is today rather than deliberately. Secondly, I do not feel it is giving enough credit to the idea that protobuf simply was not designed to be elegant in the first place. It's not a piece of modern art. It is not a Ferrari. It is a Ford F150. What protobuf does have right is the stuff that's boring; it has a ton of language bindings that are mostly pretty efficient and complete. I'm not going to say it's all perfect, because it's not. But protobuf has an ecosystem that, say, Cap'n Proto does not. This matters more than the fact that protobuf is inelegant. Protobuf may be inelegant, but it does in fact get the job done for many of the use cases. In my opinion, protobuf is actually good enough. It's not that ugly. Most programming languages you would use it from are more inelegant and incongruent than it is. I don't feel like modelling data or APIs in it is overly limited by this. Etc, etc. That said, if you want to make the most of protobuf, you'll want to make use of more than just protoc, and that's the benefit of having a rich ecosystem: there are plenty of good tools and libraries built around it. I'll admit that outside of Google, it's a bit harder to get going, but my experiences on side projects have made me feel that protobuf outside of Google is still a great idea. It is, in my opinion, significantly less of a pain in the ass than using OpenAPI based solutions or GraphQL, and much more flexible than language-specific tools like tRPC. For a selection of what the ecosystem has to offer, I think this list is a good start. https://github.com/grpc-ecosystem/awesome-grpc#protocol-buffers https://github.com/grpc-ecosystem/awesome-grpc#protocol-buff...
- hot_gril 4y ago> Protobuf is the worst serialization format and schema IDL, except for all the others. I don't even feel this way. I can complain for days about most technical things, including things Google has made (e.g. Golang), but not protobufs. They're just good. They don't totally replace JSON of course, but for the intended use case of internal distributed systems, they're solid.
- lmm 4y agoMissing the point. The whole point of protobufs is to be a stable wire format that's forward compatible. They can't have sum types because you can't add one of those in a forward compatible way; oneof works the way it does because that's the only way it can. (I do think repeated is overengineered and should just be a collection field, which would solve the problem of not composing with oneof, but that's honestly just a minor wart). Also if your language's bindings don't preserve the relationship between x and has_x and don't play nice with recursion schemes, use a better binding generator! You don't have to use the default one.
- vlmutolo 4y agoTypical is an interesting new serialization format that supports sum types. It introduces a mechanism called “asymmetric” fields that’s pretty clever. https://github.com/stepchowfun/typical https://github.com/stepchowfun/typical
- pcj-github 4y ago"They’re clearly written by amateurs, unbelievably ad-hoc, mired in gotchas, tricky to compile, and solve a problem that nobody but Google really has." Quite a statement from someone that spent less than a year at Google. Personally I really appreciate protobuf; it has saved me and teams I've worked with a ton of time and effort. Putting this article on my ignore list.
- cryptos 4y agoIf it is actually so terrible, I wonder why no player in the industry (e.g. Microsoft or some other big player) comes up with a better protocol.
- another2another 4y agoThey did, in so far as they benefited their own ecosystems. Microsoft has DCOM and other .net remoting mechanisms, which may even work over the Internet. Next/Apple has Distributed Objects: https://developer.apple.com/library/archive/documentation/Cocoa/Conceptual/DistrObjects/DistrObjects.html#//apple_ref/doc/uid/10000102i https://developer.apple.com/library/archive/documentation/Co... And databases have had it since forever, you throw in some SQL and you get back some data over defined protocols like TDS. You can even invoke some small programs via stored procs. I have no idea WTF SOAP is though …
- dekhn 4y agoI've worked with many serialization formats (XDR, SOAP XML w/ XML schema, CORBA IDL and IIOP, JSON with and without schema, pickle, and many more). Protocol buffers remind me of XDR (https://en.wikipedia.org/wiki/External_Data_Representation https://en.wikipedia.org/wiki/External_Data_Representation). Which was a great technology in the day; NFS and NIS were implemented using it. I do agree with some of the schema modelling criticisms in this article, but the ultimate thing to understand is: protocol buffers were invented to allow google to upgrade servers of different services (ads and search) asychronously and still be able to pass messages between them that can be (at least partly) decoded. They were then adopted for wide-range data modelling and gained a number of needed features, but also evolved fairly poorly. IIRC you still can't make a 4GB protocol buffer because the Java implementation required signed integers (experts correct me if I'm wrong) and wouldn't change. This is a problem when working with large serialized DL models. I was talking to Rob Pike over coffee one morning and he said everythign could be built with just nested sequences key/value pairs and no schema (IE, placing all the burden for decoding a message semantically on the client) and I don't think he's completely wrong.
- AtNightWeCode 4y agoI worked a lot with protocol buffers and for the problem it claims to solve it solves that problem well. I can't really grasp this criticism. That extreme type hell that Java coders like is not really a thing outside the Java world.