9 ms·
Protocol Buffers v3.0.0 released
- mattiemass 10y agoWow, this seems to address a bunch of problems I've experienced with protobuf in the past. Looks awesome!
- grosbisou 10y agoCould you expand on the problems you encountered?
- colanderman 10y agoI've never looked at proto3, but proto2 has at least the following issues: * No clue about namespacing. If you pick the wrong name for something, you can have name clashes within a protobuf, across uninterpreted option classes, with protobuf source code, with your own source code; and it's different if you're in Python or C. Nowhere are naming restrictions defined. * The API is maddening and inconsistent, especially in Python. (It's totally different between Python and C.) Some things look like lists but really aren't (e.g. you can't assign a list to a repeated field in Python). Even basic reflection (e.g. to get at uninterpreted options) is a Lovecraftian nightmare, and the docs are wholly unhelpful. * Good luck serializing a list. There's not really such a thing, despite that the API pretends like there is; there are only repeated fields. So you need a separate flag to distinguish "empty list" from "not present list". * Abstruse implementation. There are so many layers of indirection in the generated source and the core library that I wouldn't know where to start debugging. Not sure if they fixed any of these issues with proto3.
- cbsmith 10y agoThe short answer is the Python implementation wasn't exactly great.
- colanderman 10y agoReflection in the C++ version is as bad or worse given that you can't mess around with it in a REPL to figure out how it really works. And the C++ version has most of the namespacing issues (e.g. any field starting with "set_" has potential to clash with another field). Both implementations are equally bad, despite that they seem to have been written by two separate teams that didn't communicate with each other.
- cbsmith 10y agoHonestly, I never had much trouble with the C++ one.
- deleted 10y ago[deleted]
- mattiemass 10y agoDealing with forward- and backwards-compatibility with enum changes has bit me many times in the past. So has required fields.
- JoachimSchipper 10y agoThis looks like a nice evolution. It's a pity that the "deterministic serialization" gives so few guarantees; I have worked on at least one project that really needed this. (Basically, we wanted to parse a signed blob, do some work, and pass the original data on without breaking the signature; unfortunately, this requires keeping the serialized form around, since the serialized form cannot be re-generated from its parsed format.)
- cbsmith 10y agoIn a trusted system, if you don't trust the structure you are working with, why would you trust the signature? I'd want to always work from the signed blob. That said, this is one reason to use flatbuffers/capt'n proto I guess: you don't have to worry about this since you never unpack the blob.
- JoachimSchipper 10y agoThink of a data flow A->B->C, with A e.g. handling incoming message server, B being a spam/virus filter, and C holding the user's mailbox. Spam/virus filters are useful, but are also rather vulnerable - so C is willing to trust B's spam/non-spam judgement, but wants to ensure that B can't alter or make up messages. If protobufs had one canonical encoding, B could unpack the message and re-pack it when done; with the current protobuf implementation, B needs to keep the original blob around. In either case, C needs to check the signature on whatever blob it receives. (Some details have been changed.)
- cbsmith 10y agoSo wouldn't you stick with the original message from A, and just have B sign that? You wouldn't want to have B repack it, because then B has the potential to muck with things.
- pherl 10y agoThe main concern that the deterministic serialization isn't canonical is due to the unknown fields. As string and message type share the same wire type, when parsing an unknown string/message type, the parser has no idea whether to recursively canonicalize the unknown field. The cross-language inconsistency is mainly due to the string fields comparison performance, i.e. java/objc uses utf16 encodings which has different orderings than utf8 strings due to surrogate pairs. Feel free to start an issue on the github site asking for canonical serialization with your use case. We may change the deterministic serialization with stronger guarantee (e.g. cross language consistency) or add another API for canonical serialization.
- jalfresi 10y ago"The main intent of introducing proto3 is to clean up protobuf before pushing the language as the foundation of Google's new API platform" Does anyone know if this means Google's public APIs will be proto3 based? I quite like protobufs.
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- agency 10y agoThey've been experimenting[1] with exposing Google Cloud Platform APIs over gRPC (which is powered by proto3), so it seems quite likely. [1] https://cloud.google.com/blog/big-data/2016/03/announcing-grpc-alpha-for-google-cloud-pubsub https://cloud.google.com/blog/big-data/2016/03/announcing-gr...
- forrestthewoods 10y agoGoogle also has flatbuffers. I wonder if flatbuffers is being used by enough developers to justify significant development? https://github.com/google/flatbuffers https://github.com/google/flatbuffers
- IshKebab 10y agoI think it's more that GRPC (Google's RPC-over-HTTP2 protocol) directly supports Protobuf, and not Flatbuffers. All of Google's Cloud APIs use Protobuf (for example the [Speech API](https://cloud.google.com/speech/reference/rpc/ https://cloud.google.com/speech/reference/rpc/) ). I have to say, GRPC is pretty great. It's statically typed, supports loads of languages, the interfaces are simple to define (basically Protobuf), and it supports streaming requests! Most RPC systems omit that, or only have message streams (e.g. MQTT). Good RPC systems need both. The only downside I find is that it is rather complicated (in design; not use).
- forrestthewoods 10y agoAs an FYI, GRPC support was added to flatbuffers a month ago. https://github.com/google/flatbuffers/tree/master/grpc https://github.com/google/flatbuffers/tree/master/grpc
- alfalfasprout 10y agoBeen using flatbuffers in production for a high speed market feed for a month now. Love it. Decode/encode time is absurdly fast (~1-2 microseconds for a small to medium schema). If you're pushing 50k+ events/second it can be a great choice. Takes up almost no space on the wire too.
- zbjornson 10y ago> primitive fields set to default values (0 for numeric fields, empty for string/bytes fields) will be skipped during serialization. I don't totally understand this. Presumably during deserialization they will be set to defaults and not missing? Otherwise, coupled with the removal of required fields, it seems impossible to actually send a 0-value number or empty string, or to send a proto without a field and not have it set to 0 or "" (have to explicitly null the field?).
- prattmic 10y agoWithin the API, proto3 does not have the concept of field presence. All fields are "present" and default to their type's zero value. Since the client can handle this, there is no need to explicitly serialize default values.
- merb 10y agoand how do you send a explicit zero so that the client knows that the field is really set by the server and not the default? or a explicit empty string?
- winstonewert 10y agoIf the client really needs that information, the server must explicitly include it in a seperate field.
- tantalor 10y agoOne case where this question is important is when you are updating a record stored by the server. You only want to send fields you are changing because the record might be huge. But then how does the server distinguish between fields you didn't set and fields you want to set back to the default? The solution is to also tell the server which fields you are changing in a separate message. Example: { 'update_record': { # Set foo=bar 'foo': 'bar' }, 'fields_to_update': { 'foo': true, # Set some_int_flag=0 (default) 'some_int_flag': true } } See also "Field Masks in Update Operations" https://developers.google.com/protocol-buffers/docs/reference/csharp/class/google/protobuf/well-known-types/field-mask https://developers.google.com/protocol-buffers/docs/referenc...
- manish_gill 10y agoIf someone better informed than me can please explain - where and why would something like Protocol Buffers be useful?
- rainhacker 10y agocheck this out: http://google-opensource.blogspot.com/2008/07/protocol-buffers-googles-data.html http://google-opensource.blogspot.com/2008/07/protocol-buffe...
- dkopi 10y ago> Protocol buffers are Google's language-neutral, platform-neutral, extensible mechanism for serializing structured data – think XML, but smaller, faster, and simpler. You define how you want your data to be structured once, then you can use special generated source code to easily write and read your structured data to and from a variety of data streams and using a variety of languages. From https://developers.google.com/protocol-buffers/ https://developers.google.com/protocol-buffers/
- manish_gill 10y agoYes, I read this. It tells me what Protocol Buffers are. Faster, Smaller XML like data structures for serialisation. What are the most common use cases though? And do people only use them for performance reasons?
- skybrian 10y agoSadly the JSON format they chose isn't actually suitable for high-performance web apps. Web developers who use protobufs will continue to get by with various nonstandard JSON encodings.
- positr0n 10y agoWhy isn't is suitable? (I've never used protobufs)
- skybrian 10y agoThe fields are indexed by field names (converted to lower camel case) instead of tag numbers. It's great for readability, but it's a lot more verbose, particularly for repeated fields.
- peq 10y agoI think the only reason to use json over the binary format is readability, so why care about this?
- ambrice 10y ago> Added a new field option "json_name". By default proto field names are converted to "lowerCamelCase" in proto3 JSON format. This option can be used to override this behavior and specify a different JSON name for the field.
- skybrian 10y agoRight, but nobody's going to set that for every single protobuf field.
- ambrice 10y agoYou're right. The only people that would use it are people that a) care enough about optimization to switch out shorter tag names and b) don't care enough about optimization to switch to binary format. Probably not many..
- zellyn 10y ago- removing optional values is actually quite nice. In practice, I end up checking for "missing or empty string" anyway. - the "well-known types" boxed primitive types essentially add optional values back in. And depending on your language bindings, may look the same. - extensions are still allowed in proto3 syntax files, but only for options - since the descriptor is still proto2. It seems odd to build a proto3 that couldn't represent descriptors. - I still don't understand the removal of unknown fields. Reserialization of unknown fields was always the first defining characteristic of protobufs I described to people. I actually read many of the design/discussion docs internally when I worked at Google, and I still couldn't figure this one out. Although it's certainly simpler… - Protobufs are the "lifeblood" (Rob Pike's words) of Google: the protobuf team is working to get rid of significant Lovecraftian internal cruft, after which their ability to incorporate open source contributions should improve dramatically.
- tantalor 10y ago> removing optional values Slight correction: optional values are not removed. Quite the opposite; the "optional" keyword is removed because now all fields are optional. It is actually required values which were removed.
- zellyn 10y agoTrue. But when using them, it feels like every field is "present" and you don't have to worry about the "optional, missing" case.
- tantalor 10y agoFor primitives, yes, but not messages.
- honkhonkpants 10y agoYou're both right. What has been removed is the concept of presences altogether.
- teacup50 10y ago
- deleted 10y ago[deleted]
- wehadfun 10y agoIn C# why use Protocol Buffer over the XML or binary serializes?
- bmm6o 10y agoPerformance and data size are much better with protobufs: http://stackoverflow.com/questions/549128/fast-and-compact-object-serialization-in-net http://stackoverflow.com/questions/549128/fast-and-compact-o.... Built-in serializers are only workable when both ends are on the same platform (i.e. .Net), and even then class versioning can be a problem.
- klodolph 10y agoThe C# binary serializer is not really comparable in terms of what it does. It's more like Python's Pickle library. http://stackoverflow.com/questions/703073/what-are-the-deficiencies-of-the-built-in-binaryformatter-based-net-serializati http://stackoverflow.com/questions/703073/what-are-the-defic... C# binary serialization is only useful in certain circumstances. It doesn't work outside the .NET world and it even has compatibility problems within the .NET world—you can break deserialization by making certain changes to your code. From the Microsoft documentation: > The state of a UTF-8 or UTF-7 encoded object is not preserved if the object is serialized and deserialized using different .NET Framework versions. (From https://msdn.microsoft.com/en-us/library/72hyey7b(v=vs.110).aspx https://msdn.microsoft.com/en-us/library/72hyey7b(v=vs.110)....) Also see https://msdn.microsoft.com/en-us/library/ms229752(v=vs.110).aspx https://msdn.microsoft.com/en-us/library/ms229752(v=vs.110)....
- recursive 10y agoYour message will be about 5% the size of the xml one, and it will be backwards compatible, unlike the built-in binary serializer.
- blt 10y agoI was hoping for packed serialization of non-primitive types. I once used Protobuf to serialize small point clouds, and ended up needing to serialize them as a packed double array and reconstruct the (x, y, z) structure at read time to avoid Protobuf malloc'ing each point individually. Not a huge deal, but it would be a real pain for more complex types.
- amluto 10y agoThey added a feature that impressively fails to interoperate with the rest of the world. > Added well-known type protos (any.proto, empty.proto, timestamp.proto, duration.proto, etc.). Users can import and use these protos just like regular proto files. Additional runtime support are available for each language. From timestamp.proto: // A Timestamp represents a point in time independent of any time zone // or calendar, represented as seconds and fractions of seconds at // nanosecond resolution in UTC Epoch time. It is encoded using the // Proleptic Gregorian Calendar which extends the Gregorian calendar // backwards to year one. It is encoded assuming all minutes are 60 // seconds long, i.e. leap seconds are "smeared" so that no leap second // table is needed for interpretation. Nice, sort of -- all UTC times are representable. But you can't display the time in normal human-readable form without a leap-second table, and even their sample code is wrong is almost all cases: // struct timeval tv; // gettimeofday(&tv, NULL); // // Timestamp timestamp; // timestamp.set_seconds(tv.tv_sec); // timestamp.set_nanos(tv.tv_usec * 1000); That's only right if you run your computer in Google time. And, damn it, Google time leaked out into public NTP the last time their was a leap second, breaking all kinds of things. Sticking one's head in the sand and pretending there are no leap seconds is one thing, but designing a protocol that breaks interoperability with people who don't bury their heads in the sand is another thing entirely. Edit: fixed formatting
- madgar 10y ago> designing a protocol It's not a full protocol. It's a data type for a serialization library. You can write your own data types and they serialize just as well as the built-in types. > that breaks interoperability Wait, what was "broken" here? What was working before that isn't with this new release? What does this inclusion of a utility data type in a serialization library break that previously was intact?
- justinsaccount 10y agoIt's interesting that you refer to a huge amount of planning and engineering as "sticking your head in the sand". https://googleblog.blogspot.com/2011/09/time-technology-and-leaping-seconds.html https://googleblog.blogspot.com/2011/09/time-technology-and-... I think that the approach everything else uses is the "sticking your head in the sand approach". You basically pretend that there is no problem and that time is perfectly accurate, up until you have a minute with 59 or 61 seconds. Just because suddenly trying to handle "Oh shit, everything is off by an entire second!" is the approach everything else uses doesn't mean it is the right approach.
- rdtsc 10y agoHow does this compare or in general why would you pick this vs newer formats like Cap'n'proto or FlatBuffers? From FlatBuffers overview I see this comparison: --- Protocol Buffers is indeed relatively similar to FlatBuffers, with the primary difference being that FlatBuffers does not need a parsing/ unpacking step to a secondary representation before you can access data, often coupled with per-object memory allocation. The code is an order of magnitude bigger, too. Protocol Buffers has neither optional text import/export nor schema language features like unions. --- So are the newer ones useful mostly when serialization vs deserialization speed matters (https://google.github.io/flatbuffers/ https://google.github.io/flatbuffers/) ?
- cbsmith 10y agoAlso when you want to memory map a file/have live objects in shared memory, or in general have your in-memory & serialized structures be the same.
- jackmott 10y agoCap'n'proto is more or less abandoned I believe. But it and the flatbuffer approach gives very fast serialization and deserialization speed (essentially takes 0 times) but you pay a cost when you later access data, because it extracts the values you need on demand from the raw bytes. I'm not sure it would often make much sense overall.
- dwrensha 10y ago> Cap'n proto is more or less abandoned I believe As maintainer of capnproto-rust, I beg to differ. :) Cap'n Proto is indeed actively maintained, and here at Sandstorm we depend on it every day as a core piece of our infrastructure.
- ocdtrekkie 10y agoI would be very hesitant to call Cap'n Proto "abandoned". The Cap'n Proto developer is actively building a platform on top of it, and implements features in it as necessary, and as far as I've seen, actively works with pull requests for other features as well. https://github.com/sandstorm-io/capnproto/commits/master https://github.com/sandstorm-io/capnproto/commits/master https://github.com/sandstorm-io/capnproto/pulse/monthly https://github.com/sandstorm-io/capnproto/pulse/monthly
- andrewmcwatters 10y agoCould someone explain to me why you would use Protocol Buffers, Cap'n Proto, etc versus rolling your own type-length-value protocol besides API interop? What if your team could write a smaller TLV protocol, and it was necessary to keep your codebase small? Would this not be wise? Are Protobufs and party not comparable to TLV protocols?
- dyoo1979 10y agoThe efforts toward making the protocol robust might be helpful, depending on context. https://groups.google.com/d/topic/protobuf/DwyPEnvFJ-o/discussion https://groups.google.com/d/topic/protobuf/DwyPEnvFJ-o/discu...
- euyyn 10y agoIn the vast majority of cases, you want your team to spend their time doing something other than reinventing protos, debugging the in-house implementation, maintaining the library, etc. It's not clear to me anyway how doing it yourself would help keeping your codebase small vs using protos. In terms of code to maintain, doing it yourself is a net loss. In terms of binary size and method count, the proto libraries for Objective-C and Android are optimized like crazy.
- andrewmcwatters 10y agoThose are all reasons why I wanted to use protobufs to begin with. It sounded like it solved many issues for us. But I'm thinking about scripting environments, where the data types used in protobufs don't exist in the host language. Simple things like this. I think in the implementations I've seen, they're just coerced or ignored. That's fine, imo. But in terms of small codebases: a simple TLV protocol, where only limited data types are implemented, can be 1/10th of the size of any protobufs implementation. My team has built out a high performance type-length-value system that doesn't require compiled schemas for game development, and we have a very small serialization lib that's smaller than any protobufs implementation for our target language. I'd like to use protobufs to decrease the amount of modules we have to personally maintain, but I don't see the value in doing so for our particular situation.
- gonyea 10y agoShocking! Google's started supporting more languages than just the ones they care about. I really hope this signals the death of their disdain culture. Being a worthwhile Cloud provider means hiring experts in all sorts of languages and supporting their efforts. Imagine a world where Google didnt just "support node" (YEARS late), but actually turned their v8 expertise into a Cloud product. But that'd involve convincing Java-devs-turned-VPs to care about JavaScript, <2004>and EVERYONE knows that JavaScript is a terrible language.</2004>