17 ms·
Protobuffers Are Wrong (2018)
- taeric 1y agoI'm more than a little curious what event caused such a strong objection to protobuffers. :D I do tend to agree that they are bad. I also agree that people put a little too much credence in "came from Google." I can't bring myself to have this much anger towards it. Had to have been something that sparked this.
- mrits 1y agoI've used them almost daily for 15 years. They are way down the list of things I'd want improved. It has been interesting to see the protobuffers killers die out every few years though
- rimunroe 1y agoI'm just a frontend developer so most of my exposure is just as an API consumer and not someone working on the service side of things. That said: A few years ago I moved to a large company where protobufs were the standard way APIs were defined. When I first started working with the generated TypeScript code, I was confused as to why almost all fields on generated object types were marked as optional. I assumed it was due to the way people were choosing to define the API at first, but then I learned this was an intentional design choice on the part of protobufs. We ended up having to write our own code to parse the responses from the "helpfully" generated TypeScript client's responses. This meant we had to also handle rejecting nonsensical responses where an actually required field wasn't present, which is exactly the sort of thing I'd want generated clients to do. I would expect having to do some transformation myself, but not to that degree. The generated client was essentially useless to us, and the protocol's looseness offered no discernible benefit over any other API format I've used. I imagine some of my other complaints could be solved with better codegen tools, but I think fundamentally the looseness of the type system is a fatal issue for me.
- thinkharderdev 1y agoYeah, as soon as you have a moderately complex type the generated code is basically useless. Honestly, ~80% of my gripes about protocol buffers could be alleviated by just allowing me to mark a message field as required.
- iamdelirium 1y agoYou think you do but you really don't. What happens if you mark a field as required and then you need to delete it in the future? You can't because if someone stored that proto somewhere and is no longer seeing the field, you just broke their code.
- ozgrakkurt 1y agoMaybe you don’t delete it then?
- taeric 1y agoI mean, this is essentially the same lesson that database admins learn with nullable fields. Often it isn't the "deleting one is hard" so much as "adding one can be costly." It isn't that you can't do it. But the code side of the equation is the cheap side.
- thinkharderdev 1y agoIf you need to deserialize an old version then it's not a problem. The unknown field is just ignored during deserialization. The problem is adding a required field since some clients might be sending the old value during the rollout. But in some situations you can be pretty confident that a field will be required always. And if you turn out to be wrong then it's not a huge deal. You add the new field as optional first (with all upgraded clients setting the value) and then once that is rolled out you make it required. And if a field is in fact semantically required (like the API cannot process a request without the data in a field) then making it optional at the interface level doesn't really solve anything. The message will get deserialized but if the field is not set it's just an immediate error which doesn't seem much worse to me than a deserialization error.
- thinkharderdev 1y agoI feel like I could have written an article like this at various points. Probably while spending two hours trying to figure out a way to represent some protobuf type in a sane way internally.
- mike_hearn 1y agoHe says that in the article; he had to work on a "compiler" project that was much harder than it should have been because of protobuf's design choices.
- taeric 1y agoYeah, I saw that. I took that as something that happened in the past, though. Certainly colored a lot of the thinking, but feels like something more immediate had to have happened. :D
- jandrese 1y agoAs a developer I always see "came from Google" as a yellow flag. Too often I find something mildly interesting, but then realize that in order for me to try to use it I need to set up a personal mirror of half of Google's tech stack to even get it to start.
- ndr 1y agoNot even before the first line ends you get "They’re clearly written by amateurs". This is a rage bait, not worth the read.
- jilles 1y agoThe best way to get your point across is by starting with ad-hominem attacks to assert your superior intelligence.
- notmyjob 1y agoI disagree, unless you are in the majority.
- perching_aix 1y agoIs this in reference to the blogpost, the comment above, or your own comment? Cause it honestly works for all of them.
- sieabahlpark 1y ago[dead]
- tshaddox 1y agoIMO it's a pretty reasonable claim about experience level, not intelligence, and isn't at all an ad hominem attack because it's referring directly to the fundamental design choices of protocol buffers and thus is not at all a fallacy of irrelevance.
- compiler-guy 1y agoWhatever else Jeff Dean and Sanjay Ghemawat are, and whatever mistakes they made in designing protobufs, they are not amateurs. Not long after they designed and implemented protobuffers, they shared the ACM prize in computing, as well as many other similar honors. And the honors keep stacking up. None of this means that protobufs are perfect (or even good), but it does mean they weren't amateurs when they did it. https://en.wikipedia.org/wiki/Jeff_Dean https://en.wikipedia.org/wiki/Jeff_Dean https://en.wikipedia.org/wiki/Sanjay_Ghemawat https://en.wikipedia.org/wiki/Sanjay_Ghemawat
- bbkane 1y agoThe author makes good arguments; I wish they'd offered some alternatives. Despite issues, protobufs solve real problems and (imo) bring more value than cost to a project. In particular, I'd much rather work with protobufs and their generated ser/de than untyped json
- jeffbee 1y agoType system fans are so irritating. The author doesn't engage with the point of protocol buffers, which is that they are thin adapters between the union of things that common languages can represent with their type systems and a reasonably efficient marshaling scheme that can be compact on the wire.
- dano 1y agoIt is a 7 year old article without specifying alternatives to an "already solved problem." So HN, what are the best alternatives available today and why?
- deleted 1y ago[deleted]
- gsliepen 1y agoSomething like MessagePack or CBOR, and if you want versioning, just have a version field at the start. You don't require a schema to pack/unpack, which I personally think is a good thing.
- rapsey 1y agoThere are none, protobufs are great.
- nicce 1y agoDepends. ASN.1 is a beast and another industry standard, but unfortunately the best tooling is closed source.
- cryptonector 1y agoThere was ZERO PB tooling in 2000. Just write it for ASN.1 instead.
- 1y ago
- lalaithion 1y agoProtocol buffers suck but so does everything else. Name another serialization declaration format that both (a) defines which changes can be make backwards-compatibly, and (b) has a linter that enforces backwards compatible changes. Just with those two criteria you’re down to, like, six formats at most, of which Protocol Buffers is the most widely used. And I know the article says no one uses the backwards compatible stuff but that’s bizarre to me – setting up N clients and a server that use protocol buffers to communicate and then being able to add fields to the schema and then deploy the servers and clients in any order is way nicer than it is with some other formats that force you to babysit deployment order. The reason why protos suck is because remote procedure calls suck, and protos expose that suckage instead of trying to hide it until you trip on it. I hope the people working on protos, and other alternatives, continue to improve them, but they’re not worse than not using them today.
- jitl 1y agoNot widely used but I like Typical's approach https://github.com/stepchowfun/typical https://github.com/stepchowfun/typical > Typical offers a new solution ("asymmetric" fields) to the classic problem of how to safely add or remove fields in record types without breaking compatibility. The concept of asymmetric fields also solves the dual problem of how to preserve compatibility when adding or removing cases in sum types.
- cornstalks 1y agoI've never heard of Typical but the fact they didn't repeat protobuf's sin regarding varint encoding (or use leb128 encoding...) makes me very interested! Thank you for sharing, I'm going to have to give it a spin.
- zigzag312 1y agoIt looks similar to how vint64 lib encodes varints. Total length of varint can be determined via the first byte alone.
- 1y ago
- Analemma_ 1y agoThe "no enums as map keys" thing enrages me constantly. Every protobuf project I've ever worked with either has stringly-typed maps all over the place because of this, or has to write its own function to parse Map<String, V> into Map<K, V> from the enums and then remember to call that right after deserialization, completely defeating the purpose of autogenerated types and deserializers. Why does Google put up with this? Surely it's the same inside their codebase.
- riku_iki 1y agoAnd v1 and v2 protos didn't even have maps. Also, why you use string as a key and not int?
- Arainach 1y agoproto2 absolutely supported the map type.
- riku_iki 1y agoIt could be, it looks like there was some versions misalignment: The maps syntax is only supported starting from v3.0.0. The "proto2" in the doc is referring to the syntax version, not protobuf release version. v3.0.0 supports both proto2 syntax and proto3 syntax while v2.6.1 only supports proto2 syntax. For all users, it's recommended to use v3.0.0-beta-1 instead of v2.6.1. https://stackoverflow.com/questions/50241452/using-maps-in-protobuf-v2 https://stackoverflow.com/questions/50241452/using-maps-in-p...
- Arainach 1y agoMaps are not a good fit for a wire protocol in my experience. Different languages often have different quirks around them, and they're non-trivial to represent in a type-safe way. If a Map is truly necessary I find it better to just send a repeated Message { Key K, Value V } and then convert that to a map in the receiving end.
- dweis 1y agoI believe that the reason for this limitation is that not all languages can represent open enums cleanly to gracefully handle unknown enums upon schema skew.
- mountainriver 1y ago> Protobuffers correspond to the data you want to send over the wire, which is often related but not identical to the actual data the application would like to work with This sums up a lot of the issues I’ve seen with protobuf as well. It’s not an expressive enough language to be the core data model, yet people use it that way. In general, if you don’t have extreme network needs, then protobuf seems to cause more harm than good. I’ve watched Go teams spend months of time implementing proto based systems with little to no gain over just REST.
- nicce 1y agoOn the other hand, ASN.1 is very expressive and can cover pretty much anything, but Protobuff was created because people thought ASN.1 is too complex. I guess we can't have both.
- jandrese 1y ago"Those who cannot remember the past are condemned to repeat it" -- George Santayana
- theamk 1y agoOh, I remember ASN.1 very well, and I would not want to repeat it again. Protobufs have lots of problems, but at least they are better than ASN.1!
- cryptonector 1y agoDetails please. Things people say who know very little about ASN.1: - it's bloated! (it's not) - it's had lots of vulnerabilities! (mainly in hand-coded codecs) - it's expensive (it's not -- it's free and has been for two decades) - it's ugly (well, sure, but so is PB's IDL) - the language is context-dependent, making it harder to write a parser for (this is quite true, but so what, it's not that big a deal) The vulnerabilities were only ever in implementations, and almost entirely in cases of hand-coded codecs, and the thing that made many of these vulnerabilities possible was the use of tag-length-value encoding rules (BER/DER/CER) which, ironically, Protocol Buffers bloody is too. If you have a different objections to ASN.1, please list them.
- jt2190 1y ago(2018)
- tgma 1y agoIndeed. Discussed before at https://news.ycombinator.com/item?id=35281561 https://news.ycombinator.com/item?id=35281561 https://hn.algolia.com/?q=Protobuffers+Are+Wrong https://hn.algolia.com/?q=Protobuffers+Are+Wrong
- BugsJustFindMe 1y agoI went into this article expecting to agree with part of it. I came away agreeing with all of it. And I want to point out that Go also shares some of these catastrophic data decisions (automatic struct zero values that silently do the wrong thing by default).
- sethammons 1y agoWe got bit by a default value in a DMS task where the target column didn't exist so the data wasn't replicated and the default value was "this work needs to be done." This is not pb nor go. A sensible default of invalid state would have caught this. So would an error and crash. Either would have been better than corrupt data.
- OrangeDelonge 1y agoYou mean aws dms insterted the string literal “this work needs to be done” into your db?
- sethammons 1y agoSo, that target column was called the wrong name, meaning data intended for the column never arrived, causing the default value in the database to be used, which was an integer that mapped to "this work item needs to be processed still" which led to double processing the record post dms migration
- xmddmx 1y agoI share the author's sentiment. I hate these things. True story: trying to reverse engineer macOS Photos.app sqlite database format to extract human-readable location data from an image. I eventually figured it out, but it was: A base64 encoded Binary Plist format with one field containing a ProtoBuffer which contained another protobuffer which contained a unicode string which contained improperly encoded data (for example, U+2013 EN DASH was encoded as \342\200\223) This could have been a simple JSON string.
- fluoridation 1y agoI mean... you can nest-encode stuff in any serial format. You're not describing a problem either intrinsic or unique to Protobuf, you're just seeing the development org chart manifested into a data structure.
- xmddmx 1y agoGood points this wasn't entirely a protobuf-specific issue, so much as it was a (likely hierarchical and historical set of) bad decisions to use it at all. Using Protobuffers for a few KB of metadata, when the photo library otherwise is taking multiple GB of data, is just pennywise pound foolish. Of course, even my preference for a simple JSON string would be problematic: data in a database really should be stored properly normalized to a separate table and fields. My guess is that protobuffers did play a role here in causing this poor design. I imagine this scenario: - Photos.app wants to look up location data - the server returns structured data in a ProtoBuffer - there's no easy or reasonable way to map a protobuf to database fields (one point of TFA) - Surrender! just store the binary blob in SQLITE and let the next poor sod deal with it
- tgma 1y agoYou have to take into account the fact that iPhoto app has had many iterations. The binary plist stuff is very likely the native NSArchive "object archiving (serialization)" that is done by Obj-C libraries. They probably started using protobuf at some point later after iCloud. I suspect the unicode crap you are facing may even predate Cocoaization of the app (they probably used Carbon API). So it would make it a set of historical decisions, but I am not convinced they are necessarily bad decisions given the constraints. Each layer is likely responsible for handing edge cases in the application that you and I are not privy to.
- mkl95 1y agoIf you mostly write software with Go you'll likely enjoy working with protocol buffers. If you use the Python or Ruby wrappers you'd wish you had picked another tech.
- jonathrg 1y agoThe generated types in go are horrible to work with. You can't store instances of them anywhere, or pass them by value, because they contain a bunch of state and pointers (including a [0]sync.Mutex just to explicitly prohibit copying). So you have to pass around pointers at all times, making ownership and lifetime much more complicated than it needs to be. A message definition like this message AppLogMessage { sint32 Value1 = 1; double Value2 = 2; } becomes type Example struct { state protoimpl.MessageState xxx_hidden_Value1 int32 xxx_hidden_Value2 float64 xxx_hidden_unknownFields protoimpl.UnknownFields sizeCache protoimpl.SizeCache } For [place of work] where we use protobuf I ended up making a plugin to generate structs that don't do any of the nonsense (essentially automating Option 1 in the article): type ExamplePOD struct { Value1 int32 Value2 float64 } with converters between the two versions.
- vander_elst 1y agoAlways initializing with a default and no algebraic types is an always loaded foot gun. I wonder if the people behind golang took inspiration from this.
- wrsh07 1y agoThe simplest way to understand go is that it is a language that integrates some of Google's best cpp features (their lightweight threads and other multi threading primitives are the highlights) Beyond that it is a very simple language. But yes, 100%, for better and worse, it is deeply inspired by Google's codebase and needs
- iamdelirium 1y agoYeah, oneOf fields can be repeated but you can just wrap them in a message. It's not as pretty but I've never had any issues with this. The fact that the author is arguing for making all messages required means they don't understand the reasoning for why all fields are optional. This breaks systems (there are are postmortems outlining this) then there are proto mismatches .
- nu11ptr 1y agoShould have (2018) call out
- ants_everywhere 1y ago> Maintain a separate type that describes the data you actually want, and ensure that the two evolve simultaneously. I don't actually want to do this, because then you have N + 1 implementations of each data type, where N = number of programming languages touching the data, and + 1 for the proto implementation. What I personally want to do is use a language-agnostic IDL to describe the types that my programs use. Within Google you can even do things like just store them in the database. The practical alternative is to use JSON everywhere, possibly with some additional tooling to generate code from a JSON schema. JSON is IMO not as nice to work with. The fact that it's also slower probably doesn't matter to most codebases.
- thinkharderdev 1y ago> I don't actually want to do this, because then you have N + 1 implementations of each data type, where N = number of programming languages touching the data, and + 1 for the proto implementation. I think this is exactly what you end up with using protobuf. You have an IDL that describes the interface types but then protoc generates language-specific types that are horrible so you end up converting the generated types to some internal type that is easier to use. Ideally if you have an IDL that is more expressive then the code generator can create more "natural" data structures in the target language. I haven't used it a ton, but when I have used thrift the generated code has been 100x better than what protoc generates. I've been able to actually model my domain in the thrift IDL and end up with types that look like what I would have written by hand so I don't need to create a parallel set of types as a separate domain model.
- danans 1y ago> The practical alternative is to use JSON everywhere, possibly with some additional tooling to generate code from a JSON schema. Protobuf has a bidirectional JSON mapping that works reasonably well for a lot of use cases. I have used it to skip the protobuf wire format all together and just use protobuf for the IDL and multi-language binding, both of which IMO are far better than JSON-Schema. JSON-Schema is definitely more powerful though, letting you do things like field level constraints. I'd love to see you tomorrow that paired the best of both.
- MountainTheme12 1y agoI agree with the author that protobuf is bad and I ran into many of the issues mentioned. It's pretty much mandatory to add version fields to do backwards compatibility properly. Recently, however, I had the displeasure of working with FlatBuffers. It's worse.
- giveita 1y agoOut of interest why not make the version part of say the URL?
- MountainTheme12 1y agoThat one was used to implement save data in a game.
- wrsh07 1y ago> This insane list of restrictions is the result of unprincipled design choices and bolting on features after the fact I'm not very upset that protobuf evolved to be slightly more ergonomic. Bolting on features after you build the prototype is how you improve things. Unfortunately, they really did design themselves into a corner (not unlike python 2). Again, I can't be too upset. They didn't have the benefit of hindsight or other high performance libraries that we have today.
- techbrovanguard 1y agoi used protobuffers a lot at $previous_job and i agree with the entire article. i feel the author’s pain in my bones. protobuffers are so awful i can’t imagine google associating itself with such an amateur, ad hoc, ill-defined, user hostile, time wasting piece of shit. the fact that protobuffers wasn’t immediately relegated to the dustbin shows just how low the bar is for serialization formats.
- esafak 1y agoWhat do you use?
- fmbb 1y agoWell, worse is better.
- briandw 1y agoThe crappy system that everyone ends up using is better than the perfectly designed system that's only seen in academic papers. Javascript is the poster-child of Worse is Better. Protobuffs are a PITA, but they are widely used and getting new adoption in industry. https://en.wikipedia.org/wiki/Worse_is_better https://en.wikipedia.org/wiki/Worse_is_better
- BoorishBears 1y agoI worked at a company that had their own homegrown Protobuf alternative which would add friction to life constantly. Especially if you had the audacity to build anything that wasn't meant to live in the company monorepo (your Python script is now a Docker image that takes 30 minutes to build). One day I got annoyed enough to dig for the original proposal and like 99.9% of initiatives like this, it was predicated on: - building a list of existing solutions - building an overly exhaustive list, of every facet of the problem to be solved - declare that no existing solution hits every point on your inflated list - "we must build it ourselves." It's such a tired playbook, but it works so often unfortunately. The person who architects and sells it gets points for "impact", then eventually moves onto the next company. In the meantime the problem being solved evolves and grows (as products and businesses tend to), the homegrown solution no longer solves anything perfectly, and everyone is still stuck dragging along said solution, seemingly forever. - Usually eventually someone will get tired enough of the homegrown solution and rightfully question why they're dragging it along, and if you're lucky it gets replaced with something sane. If you're unlucky that person also uses it as justification to build a new in-house solution (we're built the old one after all), and you replay the loop. In the case of serialization though, that's not always doable. This company was storing petabytes (if not exabytes) of data in the format for example.
- pshirshov 1y agoI've created several IDL compilers addressing all issues of protobuf and others. This particular one provides strongest backward compatibility guarantees with automatic conversion derivation where possible: https://github.com/7mind/baboon https://github.com/7mind/baboon Protobuf is dated, it's not that hard to make better things.
- defraudbah 1y agoyou are absolutely right! what alternative do we have? sending json and base64 strings
- giveita 1y agoOr XML. Maybe C structures as stored in memory?
- defraudbah 1y agoI see your age yeah, c structures sounds about right
- allanrbo 1y agoSometimes you are integrating with system that already use proto though. I recently wrote a tiny, dependency-free, practical protobuf (proto3) encoder/decoder. For those situations where you need just a little bit of protobuf in your project, and don't want to bother with the whole proto ecosystem of codegen and deps: https://github.com/allanrbo/pb.py https://github.com/allanrbo/pb.py
- imtringued 1y ago>The solution is as follows: > * Make all fields in a message required. This makes messages product types. Meanwhile in the capnproto FAQ: >How do I make a field “required”, like in Protocol Buffers? >You don’t. You may find this surprising, but the “required” keyword in Protocol Buffers turned out to be a horrible mistake. I recommend reading the rest of the FAQ [0], but if you are in a hurry: Fixed schema based protocols like protobuffers do not let you remove fields like self describing formats such as JSON. Removing fields or switching them from required to optional is an ABI breaking change. Nobody wants to update all servers and all clients simultaneously. At that point, you would be better off defining a new API endpoint and deprecating the old one. The capnproto faq article also brings up the fact that validation should be handled on the application level rather than the ABI level. [0] https://capnproto.org/faq.html https://capnproto.org/faq.html
- ericpauley 1y agoI lost the plot here when the author argued that repeated fields should be implemented as in the pure lambda calculus... Most of the other issues in the article can be solved be wrapping things in more messages. Not great, not terrible. As with the tightly-coupled issues with Go, I'll keep waiting for a better approach any decade now. In the meantime, both tools (for their glaring imperfections) work well enough, solve real business use cases, and have a massive ecosystem moat that makes them easy to work with.
- wnoise 1y agoThey didn't. Pure lambda calculus would have been "a function that when applied to a number encoded as a function, extracts that value". They did it essentially as a linked list, C-strings, or UTF-8 characters: "current data, and is there more (next pointer, non-null byte, continuation bit set)?" They also noted that it could have this semantics without necessarily following this implementation encoding, though that seems like a dodge to me; length-prefixed array is a perfectly fine primitive to have, and shouldn't be inferred from something that can map to it.
- jsnell 1y agoDiscussed many times over the years: https://news.ycombinator.com/item?id=18188519 https://news.ycombinator.com/item?id=18188519 (299 comments) https://news.ycombinator.com/item?id=21871514 https://news.ycombinator.com/item?id=21871514 (215 comments) https://news.ycombinator.com/item?id=35281561 https://news.ycombinator.com/item?id=35281561 (59 comments)
- tptacek 1y agoThere are a lot of great comments on these old threads, and I don't think there's a lot of new science in this field since 2018, so the old threads might be a better read than today's. Here's a fun one: https://news.ycombinator.com/item?id=21873926 https://news.ycombinator.com/item?id=21873926
- gethly 1y agoI too was using PBs a lot, as they are quite popular in the Go world. But i came to the conclusion that they and gRPC are more trouble than they are worth. I switched to JSON, HTTP "REST" and websockets, if i need streaming, and am as happy as i could be. I get the api interoperability between various languages when one wants to build a client with strict schema but in reality, this is more of a theory than real life. In essence, anyone who subscribes to YAGNI understands that PB and gRPC are a big no-no. PS: if you need binary format, just use cbor or msgpack. Otherwise the beauty of json is that it human-readable and easily parseable, so even if you lack access to the original schema, you can still EASILY process the data and UNDERSTAND it as well.
- tombert 1y agoI am very partial to msgpack. It has routinely met or exceeded my performance needs and doesn’t depend on weird code generation, and is super easy to set up. Something that I don’t see talked about much with msgpack, but I think is cool: if your project doesn’t span across multiple languages, you can actually embed those language semantics into your encoder with extensions. For example, in Clojure’s port of msgpack out of the box, you can use Clojure keywords out of the box and it will parse correctly without issue. You also can have it work with sets. Obviously you could define some kind of mapping yourself and use any binary format to do this, ultimately the [en|de]coder is just using regular msgpack constructs behind the scenes, but i have always had to do that manually while with msgpack it seems like the libraries readily embrace it.
- gethly 1y agoIndeed, the support is widespread across languages. OTOH, using compression, like basic gzip, for http responses, turns the text format into binary format and with http2 or http3 there is no overhead like it would be with http1. so in the end the binary aspect of these encoders might be obsolete for this use case, as long as one uses compression.
- kentonv 1y agoPrevious discussions: * https://news.ycombinator.com/item?id=18188519 https://news.ycombinator.com/item?id=18188519 * https://hn.algolia.com/?q=%22Protobuffers+Are+Wrong%22 https://hn.algolia.com/?q=%22Protobuffers+Are+Wrong%22 I guess I'll, once again, copy/paste the comment I made when this was first posted: https://news.ycombinator.com/item?id=18190005 https://news.ycombinator.com/item?id=18190005 -------- Hello. I didn't invent Protocol Buffers, but I did write version 2 and was responsible for open sourcing it. I believe I am the author of the "manifesto" entitled "required considered harmful" mentioned in the footnote. Note that I mostly haven't touched Protobufs since I left Google in early 2013, but I have created Cap'n Proto since then, which I imagine this guy would criticize in similar ways. This article appears to be written by a programming language design theorist who, unfortunately, does not understand (or, perhaps, does not value) practical software engineering. Type theory is a lot of fun to think about, but being simple and elegant from a type theory perspective does not necessarily translate to real value in real systems. Protobuf has undoubtedly, empirically proven its real value in real systems, despite its admittedly large number of warts. The main thing that the author of this article does not seem to understand -- and, indeed, many PL theorists seem to miss -- is that the main challenge in real-world software engineering is not writing code but changing code once it is written and deployed. In general, type systems can be both helpful and harmful when it comes to changing code -- type systems are invaluable for detecting problems introduced by a change, but an overly-rigid type system can be a hindrance if it means common types of changes are difficult to make. This is especially true when it comes to protocols, because in a distributed system, you cannot update both sides of a protocol simultaneously. I have found that type theorists tend to promote "version negotiation" schemes where the two sides agree on one rigid protocol to follow, but this is extremely painful in practice: you end up needing to maintain parallel code paths, leading to ugly and hard-to-test code. Inevitably, developers are pushed towards hacks in order to avoid protocol changes, which makes things worse. I don't have time to address all the author's points, so let me choose a few that I think are representative of the misunderstanding. > Make all fields in a message required. This makes messages product types. > Promote oneof fields to instead be standalone data types. These are coproduct types. This seems to miss the point of optional fields. Optional fields are not primarily about nullability but about compatibility. Protobuf's single most important feature is the ability to add new fields over time while maintaining compatibility. This has proven -- in real practice, not in theory -- to be an extremely powerful way to allow protocol evolution. It allows developers to build new features with minimal work. Real-world practice has also shown that quite often, fields that originally seemed to be "required" turn out to be optional over time, hence the "required considered harmful" manifesto. In practice, you want to declare all fields optional to give yourself maximum flexibility for change. The author dismisses this later on: > What protobuffers are is permissive. They manage to not shit the bed when receiving messages from the past or from the future because they make absolutely no promises about what your data will look like. Everything is optional! But if you need it anyway, protobuffers will happily cook up and serve you something that typechecks, regardless of whether or not it's meaningful. In real world practice, the permissiveness of Protocol Buffers has proven to be a powerful way to allow for protocols to change over time. Maybe there's an amazing type system idea out there that would be even better, but I don't know what it is. Certainly the usual proposals I see seem like steps backwards. I'd love to be proven wrong, but not on the basis of perceived elegance and simplicity, but rather in real-world use. > oneof fields can't be repeated. (background: A "oneof" is essentially a tagged union -- a "sum type" for type theorists. A "repeated field" is an array.) Two things: 1. It's that way because the "oneof" pattern long-predates the "oneof" language construct. A "oneof" is actually syntax sugar for a bunch of "optional" fields where exactly one is expected to be filled in. Lots of protocols used this pattern before I added "oneof" to the language, and I wanted those protocols to be able to upgrade to the new construct without breaking compatibility. You might argue that this is a side-effect of a system evolving over time rather than being designed, and you'd be right. However, there is no such thing as a successful system which was designed perfectly upfront. All successful systems become successful by evolving, and thus you will always see this kind of wart in anything that works well. You should want a system that thinks about its existing users when creating new features, because once you adopt it, you'll be an existing user. 2. You actually do not want a oneof field to be repeated! Here's the problem: Say you have your repeated "oneof" representing an array of values where each value can be one of 10 different types. For a concrete example, let's say you're writing a parser and they represent tokens (number, identifier, string, operator, etc.). Now, at some point later on, you realize there's some additional piece of data you want to attach to every element. In our example, it could be that you now want to record the original source location (line and column number) where the token appeared. How do you make this change without breaking compatibility? Now you wish that you had defined your array as an array of messages, each containing a oneof, so that you could add a new field to that message. But because you didn't, you're probably stuck creating a parallel array to store your new field. That sucks. In every single case where you might want a repeated oneof, you always want to wrap it in a message (product type), and then repeat that. That's exactly what you can do with the existing design. The author's complaints about several other features have similar stories. > One possible argument here is that protobuffers will hold onto any information present in a message that they don't understand. In principle this means that it's nondestructive to route a message through an intermediary that doesn't understand this version of its schema. Surely that's a win, isn't it? > Granted, on paper it's a cool feature. But I've never once seen an application that will actually preserve that property. OK, well, I've worked on lots of systems -- across three different companies -- where this feature is essential.
- beders 1y agoIf you opt for non-human-readable wire-formats it better be because of very important reasons. Something about measuring performance and operational costs. If you need to exchange data with other systems that you don't control, a simple format like JSON is vastly superior. You are restricted to handing over tree-like structures. That is a good thing as your consumers will have no problems reading tree-like structures. It also makes it very simple for each consumer/producer to coerce this data into structs or objects as they please and that make sense to their usage of the data. You have to validate the data anyhow (you do validate data received by the outside world, do you?), so throwing in coercing is honestly the smallest of your problems. You only need to touch your data coercion if someone decides to send you data in a different shape. For tree-like structures it is simple to add new things and stay backwards compatible. Adding a spec on top of your data shapes that can potentially help consumers generate client code is a cherry on top of it and an orthogonal concern. Making as little assumptions as possible how your consumers deal with your data is a Good Thing(tm) that enabled such useful(still?) things as the WWW.
- deleted 1y ago[deleted]
- dinobones 1y agolols, the weird protobuf initialization semantics has caused so many OMGs. Even on my team it lead to various hard to debug bugs. It's a lesson most people learns the hard way after using PBs for a few months.
- deleted 1y ago[deleted]
- fsmv 1y agoI actually really strongly prefer 0 being identical to unset. If you have an unset state then you have to check if the field is unset every time you use it. Using 0 allows you to make all of your code "just work" when you pass 0 to it so you don't need to check at all. It's like how in go most structs don't have a constructor, they just use the 0 value. Also oneof is made that way so that it is backwards compatible to add a new field and make it a oneof with an existing field. Not everything needs to be pure functional programming.
- zigzag312 1y agoAmong other things, I don't like that they won't support nullable getters/setters: https://protobuf.dev/design-decisions/nullable-getters-setters/ https://protobuf.dev/design-decisions/nullable-getters-sette...
- guzik 1y agoWe thought for a long time about using protobufs in our product [1] and in the end we went with JSON-RPC 2.0 over BLE, base64 for bigger chunks. Yeah, you still need to pass sample format and decode manually. The overhead is fine tho, debugging is way easier (also pulling in all of protobuf just wasn't fun). [1] aidlab.com/aidlab-2
- shdh 1y agoI just wish protobuf had proper delta compression out of the box
- bloppe 1y agoProtobuf's main design goal is to make space-optimized binary tag-length-value encoding easy. The mentality is kinda like "who cares what the API looks like as long as it can support anything you want to do with TLV encoding and has great performance." Things like oneofs and maps are best understood as slightly different ways of creating TLV fields in a message, rather than pieces of a comprehensive modern type system. The provided types are simply the necessary and sufficient elements to model any fuller type system using TLV.
- guelo 1y agoYes but the point is that nobody outside of super big tech has a need to optimize a few bytes here and there at the expense of atrocious devx.
- yablak 1y ago"you're the worst serialization/config format I've ever heard of"
- summerlight 1y agohttps://news.ycombinator.com/item?id=18190005 https://news.ycombinator.com/item?id=18190005 Just FYI: an obligatory comment from the protobuf v2 designer. Yeah, protobuf has lots of design mistakes but this article is written by someone who does not understand the problem space. Most of the complexity of serialization comes from implementation compatibility between different timepoints. This significantly limits design space.
- thethimble 1y agoRelatedly, most of the author's concerns are solved by wrapping things in a message. > oneof fields can’t be repeated. Wrap oneof field in message which can be repeated > map fields cannot be repeated. Wrap in message which can contain repeated fields > map values cannot be other maps. Wrap map in message which can be a value Perhaps this is slightly inconvenient/un-ergonomic, but the author is positioning these things as "protos fundamentally can't do this".
- evanmoran 1y agoTo clarify. Protobuf’s simplest change is adding a field to a message so wrapping maps of maps, maps of fields, oneof fields into a message makes these play to its strengths. It feels like over engineering to turn your Inventory map of items into a Inventory message, but you will be grateful for it when you need a capacity field later.
- deleted 1y ago[deleted]
- missinglugnut 1y ago>Most of the complexity of serialization comes from implementation compatibility between different timepoints. The author talks about compatibility a fair bit, specifically the importance of distinguishing a field that wasn't set from one that was intentionally set to a default, and how protobuffs punted on this. What do you think they don't understand?
- 1y ago
- klodolph 1y agoProtobuffers suck as a core data model. My take? Use them as a serialization and interchange format, nothing more. > This puts us in the uncomfortable position of needing to choose between one of three bad alternatives: I don’t think there is a good system out there that works for both serialization and data models. I’d say it’s a mostly unsolved problem. I think I am happy with protobufs. I know that I have to fight against them contaminating the codebase—basically, your code that uses protobufs is code that directly communicates over raw RPC or directly serializes data to/from storage, and protobufs shouldn’t escape into higher-level code. But, and this is a big but, you want that anyway. You probably WANT your serialization to be able to evolve independently of your application logic, and the easy way to do that is to use different types for each. You write application logic using types that have all sorts of validation (in the "parse, don't validate" sense) and your serialization layer uses looser validation. This looser validation is nice because you often end up with e.g. buggy code getting shipped that writes invalid data, and if you have a loose serialization layer that just preseves structure (like proto or json), you at least have a good way to munge it into the right shape. Evolving serialized types has been such a massive pain at a lot of workplaces and the ad-hoc systems I've seen often get pulled into adopting some of the same design choices as protos, like "optional fields everywhere" and "unknown fields are ok". Partly it may be because a lot of ex-Google employees are inevitably hanging around on your team, but partly because some of those design tradeoffs (not ALL of them, just some of them) are really useful long-term, and if you stick around, you may come to the same conclusion. In the end I mostly want something that's a little more efficient and a little more typed than JSON, and protos fit the bill. I can put my full efforts into safety and the "correct" representation at a different layer, and yes, people will fuck it up and contaminate the code base with protos, but I can fix that or live with it.
- barrkel 1y agoI'm afraid that this is a case of someone imagining that there are Platonic ideal concepts that don't evolve over time, that programs are perfectible. But people are not immortal and everything is always changing. I almost burst out in laughter when the article argued that you should reuse types in preference to inlining definitions. If you've ever felt the pain of needing to split something up, you would not be so eager to reuse. In a codebase with a single process, it's pretty trivial to refactor to split things apart; you can make one CL and be done. In a system with persistence and distribution, it's a lot more awkward. That whole meaning of data vs representation thing. There's fundamentally a truth in the correspondence. As a program evolves, its understanding of its domain increases, and the fidelity of its internal representations increase too, by becoming more specific, more differentiated, more nuanced. But the old data doesn't go away. You don't get to fill in detail for data that was gathered in older times. Sometimes, the referents don't even exist any more. Everything is optional; what was one field may become two fields in the future, with split responsibilities, increased fidelity to the domain.
- deleted 1y ago[deleted]
- BinaryIgor 1y agoAvro (and others) has its own set of problems as well. For messaging, JSON, used in the same way and with the same versioning practices as we have established for evolving schemas in REST APIs, has never failed me. It seems to me that all these rigid type systems for remote procedure calls introduce more problems that they really solve and bring unnecessary complexity. Sure, there are tradeoffs with flexible JSONs - but simplicity of it beats the potential advantages we get from systems like Avro or ProtoBuf.
- nice_byte 1y ago> Make all fields in a message required. funnily enough, this line alone reveals the author to be an amateur in the problem space they are writing so confidently about.
- nice_byte 1y agothe complaints about the Protobuf type system being not flexible enough are also really funny to read. fundamentally, the author refuses to contend with the fact that the context in which Protobufs are used -- millions of messages strewn around random databases and files, read and written by software using different versions of libraries -- is NOT the same scenario where you get to design your types once and then EVERYTHING that ever touches those types is forced through a type checker. again, this betrays a certain degree of amateurishness on the author's part. Kenton has already provided a good explanation here: https://news.ycombinator.com/item?id=45140590 https://news.ycombinator.com/item?id=45140590
- instig007 1y ago> is NOT the same scenario where you get to design your types once and then EVERYTHING that ever touches those types is forced through a type checker. the author never claimed the types had to be designed only once, he claimed that schema evolution chosen by protobuf is inadequate for the purpose of lossless evolution. > Kenton has already provided a good explanation here: https://news.ycombinator.com/item?id=45140590 https://news.ycombinator.com/item?id=45140590 TLDR: yada-yada [...] protobuf is practical, type algebra either doesn't exist or impractical because only PL theorists know about it, not Kenton.
- kentonv 1y ago> type algebra either doesn't exist or impractical because only PL theorists know about it, not Kenton. Hi I'm Kenton. I, too, was enamored with advanced PL theory in college. Designed and implemented my own purely-functional programming language. Still wish someone would figure out a working version of dependent types for real-world use, mainly so we could prove array bounds-safety without runtime checks. In two decades building real-world complex systems, though, I've found that getting PL theory right is rarely the highest-leverage way to address the real problems of software engineering.
- m463 1y agopersuasive or pervasive?
- bithive123 1y agoI don't know if the author is right or wrong; I've never dealt with protobufs professionally. But I recently implemented them for a hobby project and it was kind of a game-changer. At some stage with every ESP or Arduino project, I want to send and receive data, i.e. telemetry and control messages. A lot of people use ad-hoc protocols or HTTP/JSON, but I decided to try the nanopb library. I ended up with a relatively neat solution that just uses UDP packets. For my purposes a single packet has plenty of space, and I can easily extend this approach in the future. I know I'm not the first person to do this but I'll probably keep using protobufs until something better comes along, because the ecosystem exists and I can focus on the stuff I consider to be fun.
- _zoltan_ 1y agoand since it's UDP, if it's lost it's lost. and since it's not standard http/JSON, nobody will have a clue in a year and can't decode it. to learn and play with it it's fine, else why complicate life?
- Farmadupe 1y agoUsing protobuf is practical enough in embedded. This person isn't the first and won't be the last. Way faster than JSON, way slower than C structs. However protobuf is ridiculously interchangeable and there are serializers for every language. So you can get your interfaces fleshed out early in a project without having to worry that someone will have a hard time ingesting it later on. Yes it's a pain how an empty array is a valid instance of every message type, but at least the fields that you remember to send are strongly typed. And field optionality gives you a fighting chance that your software can still speak to the unit that hasn't been updated in the field for the last five years. On the embedded side, nanopb has worked well for us. I'm not missing having to hand maintain ad-hoc command parsers on the embedded side, nor working around quirks and bugs of those parsers on the desktop side
- tliltocatl 1y agoEmbedded/constrained UDP is where protobuf wire format (but not google's libraries) rocks: IoT over cellular and such, where you need to fit everything into a single datagram (number of roundtrips is what determines power consumption). As to those who say "UDP is unreliable" - what you do is you implement ARQ on the application level. Just like TCP does it, except you don't have to waste roundtrips on SYN-SYN-ACK handshake nor waste bytes on sending data that are no longer relevant. Varints for the win. Send time series as columns of varint arrays - delta or RLL compression becomes quite straightforward. And as a bonus I can just implement new fields in the device and deploy right away - the server-side support can wait until we actually need it. No, flatbuffers/cap'n'proto are unacceptably big because of fixed layout. No, CBOR is an absolute no go - why on earth would you waste precious bytes on schema every time? No, general-purpose compression like gzip wouldn't do much on such a small size, it will probably make things worse. Yes, ASN is supposed to be the right solution - but there is no full-featured implementation that doesn't cost $$$$ and the whole thing is just too damn bloated. Kinda fun that it sucks for what it is supposed to do, but actually shines elsewhere.
- ryukoposting 1y agoProtobuf's original sin was failing to distinguish zero/false from undefined/unset/nil. Confusion around the semantics of a zero value are the root of most proto-related bugs I've come across. At the same time, that very characteristic of protobuf makes its on-wire form really efficient in a lot of cases. Nearly every other complaint is solved by wrapping things in messages (sorry, product types). Don't get the enum limitation on map keys, that complaint is fair. Protobuf eliminates truckloads of stupid serialization/deserialization code that, in my embedded world, almost always has to be hand-written otherwise. If there was a tool that automatically spat out matching C, Kotlin, and Swift parsers from CDDL, I'd certainly give it a shot.
- mdhb 1y agoAgreed the CDDL to codegen pipeline / tooling is the biggest thing holding back CBOR at the moment. Some solutions do exist like here’s a C one[1] which maybe you could throw in some WASI / WASM compilation and get “somewhat” idiomatic bindings in a bunch of languages. Here’s another for Rust [2] but I’m sure I’ve seen a bunch of others around. I think what’s missing is a unified protoc style binary with language specific plugins. [1] https://github.com/NordicSemiconductor/zcbor https://github.com/NordicSemiconductor/zcbor [2] https://github.com/dcSpark/cddl-codegen https://github.com/dcSpark/cddl-codegen
- jsnell 1y ago> Protobuf's original sin was failing to distinguish zero/false from undefined/unset/nil. It's only proto3 that doesn't distinguish between zero and unset by default. Both the earlier and later versions support it. Proto3 was a giant pile of poop in most respects, including removing support for field presence. They eventually put it back in as a per-field opt-in property, but by then the damage was done. A huge unforced mistake, but I don't think a change made after the library had existed for 15 years and reverted qualifies as an "original sin".
- rednafi 1y agoI like the problems that Protobuf solves, just not the way it solves them. Protobuf as a language feels clunky. The “type before identifier” syntax looks ancient and Java-esque. The tools are clunky too. protoc is full of gotchas, and for something as simple as validation, you need to add a zillion plugins and memorize their invocation flags. From tooling to workflow to generated code, it’s full of Google-isms and can be awkward to use at times. That said, the serialization format is solid, and the backward-compatibility paradigms are genuinely useful. Buf adds some niceties to the tooling and makes it more tolerable. There’s nothing else that solves all the problems Protobuf solves.
- kiitos 1y agoalmost the entire purpose of anything like protocol buffers is to provide a safe mechanism for backwards-compatible forward changes -- "no one uses that stuff"?? what a weird and broken take
- spectraldrift 1y agoI'm not sure why this post gets boosted every few years- and unfortunately (as many have pointed out) the author demonstrates here that they do not understand distributed system design, nor how to use protocol buffers. I have found them to be one of the most useful tools in modern software development when used correctly. Not only are they much faster than JSON, they prevent the inevitable redefinition of nearly identical code across a large number of repos (which is what i've seen in 95% of corporate codebases that eschew tooling such as this). Sure, there are alternatives to protocol buffers, but I have not seen them gain widespread adoption yet.
- sylware 1y agoI don't recall properly (because I did selve my mapping projects for the moment), but don't openstreet map core data distribution format based on protobuffers?
- xyzzyz 1y agoGranted, on paper it’s a cool feature. But I’ve never once seen an application that will actually preserve that property. Chances are, the author literally used software that does it as he wrote these words. This feature is critical to how Chrome Sync works. You wouldn’t want to lose synced state if you use an older browser version on another device that doesn’t recognize the unknown fields and silently drops them. This is so important that at some point Chrome literally forked protobuf library so that unknown fields are preserved even if you are using protobuf lite mode.
- toolslive 1y agoThe author is right, but it could have been worse too. At least they were not using JSON for serialization.
- stinkbeetle 1y agoWith these serialization libraries, do any of them have a facility that allows you to specify a wire format and an application format, with recipes for converting one to the other? I haven't used these very seriously but a problem I had a while back was that that the wire format was not what the applications wanted to use, but a good application format was to space-inefficient for wire. As far as I could see there was not a great way to do this. You could rewrite wire<->app converter in every app, or have a converter program and now you essentially have two wire formats and need to put this extra program and data movement into workflows, or write a library and maintain bindings for all your languages.
- wffurr 1y ago>> You could rewrite wire<->app converter in every app This is what Google does. We joke that our entire jobs are "convert protobuf A into protobuf B".
- jzwinck 1y agoIf you care about network bandwidth you can compress before sending, as virtually all web applications do. Then you don't need to worry much about the space efficiency of the application format.
- stinkbeetle 1y agoOf the wire format you mean? I compress it and still need to care about the space efficiency of the wire format beyond that. Compression ratio does improve a lot when not doing our own, end result is significantly larger. Also it becomes also significantly slower because more data to process which is possibly the bigger problem. It's probably not like most web application, it's hardware data loggers that produce about hundreds of millions to billions of events per second (each with minimum about 4 bytes of wire format and maximum roughly 500 bytes).
- frumiousirc 1y agoThe way to do this starts with not hard-wiring the code generation step. Instead, make codegen a function of BOTH a data schema object and a code template (eg expressed in Jinja2 template language - or ZeroMQ GSL where I first saw this approach). The codegen stage is then simply the application of the template to the data schema to produce a code artifact. The templates are written assuming the data schema is provided following a meta-schema (eg JSON Schema for a data schema in JSON). One can develop, eg per-language templates to produce serialization code or intra-language converters between serialization forms (on wire) and application friendly forms. The extra effort to develop a template for a particular target is amortized as it will work across all data schemas that adhere to a common meta-schema. The "codegen" stage can of course be given non "code" templates to produce, eg, reference documentation about the data schema in different formats like HTML, text, nroff/man, etc.
- smittywerben 1y agoGoogle claimed Protobuffers are the solution but Google's planetary engineers clearly have ZERO respect for the mixed-endian remote systems keeping the galactic federation afloat with their cheap CORBA knockoff. It's like, sure which Plan 9 mainframe do you want to connect to like we all live on planet Google. Like hello???
- xg15 1y agoI'm starting to wonder if some of those bad design decisions are symptoms of a larger "cultural bias" at Google. Specifically the "No Compositionality" point: It reminds me of similar bad designs in Go, CSS and the web platform at large. The pattern seems to be that generalized, user-composable solutions are discouraged in favor of a myriad of special constructs that satisfy whatever concrete use cases seem relevant for the designers in the moment. This works for a while and reduces the complexity of the language upfront, while delivering results - but over time, the designs devolve into a rats's nest of hyperspecific design features with awkward and unintuitive restrictions. Eventually, the designers might give up and add more general constructs to the language - but those feel tacked on and have to coexist with specific features that can't be removed anymore.
- senorrib 1y agoIt works both ways. General constructs tend to become overly abstract and you end up with sneaky errors in different places due to a minor change to an abstraction. Like the old adage, this is just a matter of preference. Good software engineering requires, first and foremost, great discipline, regardless of the path or tool you choose.
- gettingoverit 1y agoIf there are errors in implementation of general constructs, they tend to be visible at their every use, and get rapidly fixed. Some general constructs are better than the others, because they have an algebraic theory behind them, and sometimes that theory was already researched for a few hundred years. For example, product/coproduct types mentioned in the article are quite close to addition and multiplication that we've all learned in school, and obey the same laws. So there are several levels where the choice of ad-hoc constructs is wrong, and in the end the only valid reason to choose them is time constraints. If they had 24 years to figure out how to do it properly, but they didn't, the technology is just dead.
- sdenton4 1y agoHm, that's idealistic... I've certainly run into cases where small changes in general systems led to hard-to-detect bugs, which took a great deal of investigation to figure out. Not all failures are catastrophic. The technology is quite alive, which is why it hasn't been 'fixed' - changing the wheels on a moving car, and all that. The actual disappointment is that a better alternative hasn't taken off in the six years since this post was written... If its so easy, where's the alternatives?
- cenamus 1y agoI really liked the typography/layout of the page, reminds me of gwern.net. But people will probably complain about serif fonts regardless
- DoneWithAllThat 1y agoThe ignorance on display in this post is truly breathtaking. I’m impressed something can be this wrong.
- co_dh 1y agoBut why do you need serialization? Because the data structure on disk is not the same as in memory. Arthur Whitney's k/q/kdb+ solved this problem by making them the same. An array has the same format in memory and on disk, so there is no serialization, and even better, you can mmap files into memory, so you don't need cache! He also removed the capability to define a structure, and force you to use dictionary(structure) of array, instead of array of structure.
- deleted 1y ago[deleted]
- throwaway127482 1y ago> But why do you need serialization? Because the data structure on disk is not the same as in memory. Not always - in browser applications for example, there is no way to directly access the disk, nevermind mmap().
- RossBencina 1y agoForget on-disk. Different CPUs represent basic data types with different in-memory representations (endianness). Furthermore different CPUs have different capabilities with respect to how data must be aligned in memory in order to read or write it (aligned/unaligned access). At least historically unaligned access could fault your process. Then there's the problem, that you allude to, that different programming languages use different data layouts (or often a non-standardised layout). If you want communication within a system comprising heterogeneous CPUs and/or languages, you need to translate or standardise your a wire format and/or provide a translation layer aka serialisation.
- sgammon 1y ago> Your guess is as good as mine for why an enum can’t be used as a map key. I filed an issue requesting this and it was denied with an explanation: https://github.com/protocolbuffers/protobuf/issues/7791#issuecomment-940507054 https://github.com/protocolbuffers/protobuf/issues/7791#issu...
- sgammon 1y ago> It’s impossible to differentiate a field that was missing in a protobuffer from one that was assigned to the default value. This is purportedly fixed in proto3 and latest SDK copies (IIRC)
- sgammon 1y ago> Contrast this behavior against message types. While scalar fields are dumb, the behavior for message fields is outright insane. The reason messages are initialized is that you can easily set a deep property path: ``` message SomeY { string example = 1; } message SomeX { SomeY y = 1; } ``` later, in java: ``` SomeX some = SomeX.newBuilder(); some.getY().setExample("hello"); // does not produce npe ``` in kotlin this syntax makes even more sense: ``` some { y.example = "hello". // does not produce npe } ```
- cryptonector 1y agoI've written several screeds in the comments here on HN about protobufs being terrible over the past few years. Basically the creators of PB ignored ASN.1 and built a bad version of mid-1980s ASN.1 and DER. Tag-length-value (TLV) encodings are just overly verbose for no good reason. They are _NOT_ "self-describing", and one does not need everything tagged to support extensibility. Even where one does need tags, tag assignments can be fully automatic and need not be exposed to the module designer. Anyone with a modicum of time spent researching how ASN.1 handles extensibility with non-TLV encoding rules knows these things. The entire arc of ASN.1's evolution over two plus decades was all about extensibility and non-TLV encoding rules! And yes, ASN.1 started with the same premise as PB, but 40 years ago. Thus it's terribly egregious that PB's designers did not learn any lessons at all from ASN.1! Near as I can tell PB's designers thought they knew about encodings, but didn't, and near as I can tell they refused to look at ASN.1 and such because of the lack of tooling for ASN.1, but of course there was even less tooling for PB since it hadn't existed. It's all exasperating.
- BobbyTables2 1y agoEven the low level implementation of protobuffers is pretty uninspiring. Adds a lot of space overhead, specially for structs only used one yet not self descriptive either. Doesn’t solve a lot of problems related to changes either. Quite frankly, too many are using up in it because it came from Google and is supposed to be some sort of divinely inspired thing. JSON, ASN.1, and even rigid C structs start to look a lot better.
- lukaslalinsky 1y agoI recently made a realization, that I can use MessagePack with a static schema defined in the code, and even pre-defined numeric field IDs, essentially replacing Protobuf for my use cases. I saw MessagePack as an alternative for JSON, with loose message structure, but it's actually a nice binary format and can be used more effectively than that. So now I enjoy things like tagged unions (in Zig/Python), and other types that are awkward to express in Protobuf. I settled on single character field names, for compatibility with msgspec, and I'm pretty happy with it. Still super compact messages with predictable schema, that are fast to parse, because I know which fields to expect.
- TeeMassive 1y agoI get the author's points, and they all are valid, but I don't understand why people would use generated code and its types through their entire project. This is just asking for trouble when the API will inevitably break as all APIs will do eventually. In our projects I mandated and pushed really hard that we create intermediary data classes that correspond one to one to the protobufs (at first). I got a lot of angry faces and reactions in PR due to the seemingly useless boiler plate code required but it saved our butts so many times when the API changed just before a release that it became the de facto standard. Also, protobufs and GRPCs are a de facto standards. Are there better alternatives? Yes. Should you use those? Most likely not because the point of serialization frameworks is to be used by many people in various tech stacks.