5 ms·
I really wish someone would create a fork/variant of Protocol Buffers for all the folks that are not afraid of versioning message schemas. With actual guarantee
by fuzzy2 4y ago
I really wish someone would create a fork/variant of Protocol Buffers for all the folks that are not afraid of versioning message schemas. With actual guarantees as to what is in the message and what isn’t.
I recently had to select a suitable data serialization format for a sort-of-distributed application. The goal was to force teams to declare/discuss contracts up-front, so schema-based. It was a very frustrating experience. All the formats with expressive schemas and type systems somehow don’t have a concept of required fields, or not anymore. How can you create a robust system like this?
- tantalor 4y agoProto2 has required. > required: a well-formed message must have exactly one of this field. https://developers.google.com/protocol-buffers/docs/proto#specifying-rules https://developers.google.com/protocol-buffers/docs/proto#sp...
- c-cube 4y agoProto3 gained "required" back relatively recently. It's even in use in Opentelemetry proto files.
- IceWreck 4y agoThrift (both fbthrift and apache thrift) has required fields
- fuzzy2 4y agoYou’re right! Dunno why I dismissed Thrift in my research, but I definitely missed this fact.
- bunderbunder 4y agoProtocol buffers used to have required fields. I worked at a company that kept using an old version specifically to keep having them. I’ve also talked to Googlers about this. My understanding is that having required fields created a robustness problem for them. My guess is that it’s the same situation as for basically everything else: the problems Google is trying to solve for, and the constraints on the solutions to those problems, are just different from what the rest of us are dealing with. That doesn’t make protocol buffers bad, but it maybe does make the level of mindshare Google gets in this corner of the technology world bad. That said, 100% agreed. Proto2 solves for that one use case (required fields), but not others such as "I want to talk to this API from the browser." I would love to have a solution that prioritized being easy to implement and support over being as heavily engineered - and, consequently, difficult to understand and use effectively - as FAANG-scale technologies tend to be.
- yaacov 4y agoRequirements change. I’ve seen 10+ year old proto files at google still being used by services that are still being actively developed. Over those ten years your data model has evolved. Your first client, the one whose data model you were imitating when you added the required field, was deprecated five years ago and turned down last year. You have a dozen or a thousand client services, each using you as a backend in a slightly different way. Are you sure every single one of them is going to require that field? Are you sure the field will even be semantically meaningful for their use case?
- bunderbunder 4y agoThose are exactly the kinds of things I was thinking of when I suggested that proto is designed for Google scale. It’s not that smaller companies never have these problems. It’s that, at smaller companies, the ways in which they manifest themselves and the cost/benefit ratios tend to favor different solutions to these problems. For example, the company I was at that stuck with Proto2 so they could keep required fields, all the engineers worked in a single room, and could resolve questions about the needs of all a protocols consumers by simply standing up and saying, "Hey everybody, …"
- xyzzy_plugh 4y agoMy understanding is that the powers that be within Google have decided that validating messages is outside the scope of schemas and serialization. protoc-gen-validate provides a portable way to perform validation: https://github.com/bufbuild/protoc-gen-validate https://github.com/bufbuild/protoc-gen-validate The problem with required fields is it kicks the can down the road when you want to deprecate a field. Keeping everything optional is much, much better for everyone in the long run.
- fuzzy2 4y agoThe problem is with default values: they are not sent over the wire. You cannot determine whether that false was deliberate or the sender just forgot to set it to true. I understand that "elastic" contracts may make some stuff easier. They do not help in forcing developers to create a message correctly, unfortunately. Still, it's great to see someone is tackling the validation rule topic. One of my stakeholders is very… enthusiastic about validation. Just goes to show how bad software engineering is in practice in this org.
- tantalor 4y agoEr no, that's not true. If you set the value of a field, then it will be serialized with that value. It doesn't matter if the value is the default for that field.
- fuzzy2 4y agoNo, it is, at least for C# and the default/"official" code generator. The docs (proto3) say this: “Also note that if a scalar message field is set to its default, the value will not be serialized on the wire.”
- tantalor 4y agoThat's true for "singular" fields, but not "optional". https://developers.google.com/protocol-buffers/docs/proto3#specifying_field_rules https://developers.google.com/protocol-buffers/docs/proto3#s... If you don't like that, don't use "singular".
- haberman 4y agoValidation is important, but the serialization layer is the wrong place to put validation logic. A protobuf is a low-level abstraction, like a struct or record type in your favorite programming language. You want validation logic to be a separate layer on top. You don't want it so coupled to parsing/serialization that you literally cannot parse/serialize something that doesn't validate. Protobuf has a rich facility for adding custom annotations to anything (messages, fields, etc). These annotations are the right place to put validation predicates. That will let anyone build a validation layer on top of protobuf, for example: https://scalapb.github.io/docs/validation/ https://scalapb.github.io/docs/validation/
- fuzzy2 4y agoIn general, I agree. The problem is that serialization (at least with proto3) is “lossy”. As I mentioned elsewhere in the discussion, proto3 messages discard certain information in the name of efficiency. The end result (message semantically invalid) does not change, of course, but the “why” could.
- haberman 4y agoYes, proto3 "singular" fields are a big problem. But now that proto3 supports "optional" fields (which remember the difference between unset and explicit 0) you can use true optional fields for all new messages going forward: https://developers.google.com/protocol-buffers/docs/proto3#specifying_field_rules https://developers.google.com/protocol-buffers/docs/proto3#s...