6 ms·
Protobuf-py: Protobuf for Python, without compromises
- fernando-ram 3mo ago[flagged]
- tuwtuwtuwtuw 3mo agoI don't know if it's a rendering issue, but (on my Android phone with Chrome), that is probably the worst readme I encountered in my life. Edit: I see now. You are spamming links to your owm site.
- usrnm 3mo agoAfter gogoproto I'm hesitant to depend on another non-standard implementation, getting off gogo was a pain. This thing may be better than the one from Google (gogo definitely was), but can we be sure that it will still be around in 10 years?
- didip 3mo agoWhat was the backstory of gogoproto?
- usrnm 3mo agoThe golang implementation of protobuf sucked historically (still does, but is improving) and gogo was an alternative that fixed a lot of problems and was nicer overall. Until its creator burned out and deprecated it. Chasing a constantly moving target that you have no control over is very taxing in the long run.
- Intralexical 3mo agoThis is concerning to hear. What do Protobufs accomplish, that requires them to be a constantly moving target?
- arccy 3mo agohttps://www.youtube.com/watch?v=HTIltI0NuNg https://www.youtube.com/watch?v=HTIltI0NuNg I think it's less that protobuf is a moving target, and more that gogo tried to add in all the features that google didn't want to maintain, and learned that maintaining a massive feature matrix was impossible.
- usrnm 3mo agoIf I remember correctly, what finally broke the camel's back was the new API that Google introduced. But I may be wrong
- esrauch 3mo agoGoGo was not a completely separate implementation but deeply hooked into the official GoProto implementation. So it wasn't "Protobuf the binary wire format" or "Protobuf the schema language" which changed over time here, changes to the Google's Go library caused it problems. It's like building a library that integrates with Jackson (a JSON library) and Jackson details changed in ways that added toil, versus JSON changing.
- crabbone 3mo agoHey. I wrote another Python implementation of Protobuf. (protopy https://gitlab.com/doodles-archive/protopy https://gitlab.com/doodles-archive/protopy it was a while ago and haven't touched it since). I'm not saying it's better than whatever this is or that it's any good, I just post it as a proof of sorts that I'm familiar with the problem. So, without further ado: Protobuf isn't a standard. You can't have a non-standard implementation of something that doesn't have a standard to begin with. In reality, you have Google's implementation for C++ and then everything else. Everything else was, for the most part, not written by Google. And it doesn't always align 100% with the C++ Google's stuff. Furthermore, C++ implementation has a lot of idiosyncrasies specific to that language that can't be translated one-to-one into other languages, or, in some cases, shouldn't be, even if they could (eg. C++ implementation is all about source code generation because generating runtime entities s.a. classes in C++ is very difficult, while in languages like Python, generating classes at runtime is easy.) Furthermore, C++ implementation has a specific way of parsing the binary payload (lazy: only the top definitions are parsed, the inner structure of messages is parsed on-demand). But, is this how every parser should behave? What if you want a SAX-like parser? ---- In the hindsight, I just think that Protobuf is not a good format for writing reliable software that aims for decades of usage. We, as in the whole programming world, don't have good formats in general, and whenever we come to the point of having to use some, we either go with an existing popular but crooked or roll our own, probably also crooked. The standard you alluded to would've been great (perhaps a refinement of ASN with more attention to parser implementation, more concrete versions etc.?) But we aren't there yet, and there isn't even a work group to try and address the issue.
- squirrellous 3mo agoHonest question - why isn’t the following document [1] a standard? Is it too loosely specified? [1] https://protobuf.dev/programming-guides/encoding/ https://protobuf.dev/programming-guides/encoding/
- 7bit 3mo agoMaybe He means because it diesen have an accepted rfc
- newswangerd 3mo agoBuf is well established and maintains a lot of protobuf packages for many languages, including the YAML implementation for go.
- lyu07282 3mo agoGoing strong since 2019! It's a wonderful company entirely dedicated to making google's miserable protobuf "somewhat" useable.
- igetspam 3mo agoCan you be sure Google’s ${anything} will be around in 6 months? They have a decades long habit of dumping things that have major use but aren’t novel internally. Worrying about the next 10 years isn’t something you can honestly do when you your argument includes Google owning a thing. Bias: I fought for things like Reader from the inside. If it doesn’t move needles, it goes away.
- est 3mo agoHope it can auto build a python class from gRPC gateway reflections.
- newswangerd 3mo agoThis is incredible news! I’ve used protobuf in Python, Go, Kotlin and Dart and the Python implementation is totally unusable. I don’t know what black magic Google uses for the Python implementation, but the classes it generates are totally opaque and impossible to inspect. I’ve been waiting for a proper python implementation for years now!
- masklinn 3mo ago> I don’t know what black magic Google uses for the Python implementation, but the classes it generates are totally opaque and impossible to inspect. TFA seems to say that they’re just thin proxies over the underlying C++ APIs, which would more than do it, and does not surprise me (the re2 Python bindings are similar, not as bad since they don’t generate Python code but they’re really c++-y — in Google’s flavour too — and uncomfortable).
- functional_dev 3mo agoI did not know this before... Google protobuf is not one thing, it is three! Same import, but: * old C++ extension * upb * pure Python upb parses FAST, but then every access is still C->Python and it slows it down. So for many reads the slow python one can win? This one helped me to dig deeper - https://vectree.io/c/how-python-protobuf-runtimes-work-pure-python-vs-upb-vs-the-c-backend https://vectree.io/c/how-python-protobuf-runtimes-work-pure-...
- masklinn 3mo ago> upb parses FAST, but then every access is still C->Python and it slows it down. So for many reads the slow python one can win? That is pretty unlikely. TFA's version is in Rust and just barely edges out upb.
- Chu4eeno 3mo agorust code tends to be incredibly slow unless you spend a ton of time optimizing (compared to other compiled languages, and contrary to popular belief).
- giovannibonetti 3mo agoI wish LaunchDarkly and other feature flag providers supported protocol buffers to MN define the feature flag schema. It would be a game changer when you have complex variations and end up reaching for untyped JSON.
- deleted 3mo ago[deleted]
- esrauch 3mo agoEngineer who works on Google Protobuf here, commenting as myself and not as an official statement. It's great to have a healthy ecosystem in the world around Protobuf. Google can't possibly fill all use cases, there's many tools and Buf makes good tooling. The Protobuf team at Google intentionally tries to enable an ecosystem around Protobuf including examples like this. Google Cloud APIs are intentionally usable with any compatible thing that can understand Protobuf encoding including this one. Kudos to Buf for making something that I'm sure a ton of people will find useful and which takes conformance so seriously. Just to chime in with some context about Google's own implementations here though (since that's a lot of the discussion otherwise). Google definitely takes Protobuf seriously including for the long term: you can't really understand how engrained it is within the Google stack without seeing it for yourself. It's not just RPC layer, it's storage, logging, FFI. Html templating is driven off Protobuf messages. Systems which interact with bank XML based systems uses Protobuf schemas. Internally it's widely used for in-memory library api types even without any direct/obvious connection to serialization just because it makes internal details like logging easier. This extremely large surface does create constraints and use-cases to balance. You can see Buf's reported numbers reflect that it is faster for a usecase they expect is typical, but at scale users do fall into the other buckets shown, affecting the performance of preexisting code is a major concern for our implementations that a greenfield implementation doesn't have. Wide exposure in critical paths alongside long term support directly causes some quirks: for example some of our APIs followed PEP8 when it was created but PEP8 changed. It looks stupid that we have wrong style APIs but also it would be stupider to break compatibility for style reasons. JavaProto as another example still supports Java8 and the runtime is compatible with 2014 gencode which is a pretty major constraint. Google Py Proto implementation has one extra interesting choice of the same gencode is reused with 3 different implementations (upb, a complete pure python one, and one that uses C++Proto as the in memory representation which libraries like TensorFlow can use to share memory between Python and C++), which is why design the way that it is with runtime created classes, the pyi files are readable but the .py files not. This definitely has pros and cons, and the direct approach taken by Buf here really makes a ton of sense. It's just that Google's maintained implementation falls into a different spot in a larger technical tradeoff space. If you see things that appear to make no sense with the official implementations, feel free to file an issue on GitHub and we can look, sometimes there is no reason and we can fix it, and sometimes there's a reason which we can explain. Kudos again to Buf here, I'm fully sure this will solve some set of real business needs better than Google's (but not because Google isn't maintaining our offerings too).
- tommek4077 3mo agoOptimizing for the last microsecond and then use Python? Why?
- dofm 3mo agoThis is a portable format that is used for interchange. The performance may matter elsewhere in your system and yet you may still want to create/access protobuf-packaged data?
- paulddraper 3mo ago> Optimizing for the last microsecond I don't see that anywhere. I see "Fast where it counts."
- quietbritishjim 3mo agoAre the only two options SIMD in hand-rolled assembler and unbearably slow? Often Python, with the most critical parts written in a compiled extension module (like this one), offers acceptable performance with enormously less complexity than writing the whole thing in a compiled language.
- sankalpmukim 3mo agoThis. Yes. This is why I clicked on the comments. Exactly
- this_was_posted 3mo agoSlightly off topic, but is anyone aware of a compile free method to convert protobuf messages to json representation based on a provided .proto file (and the other way around)? All protobuf implementations seem to require a compilation step, which makes it hard to support en-/decoding untrusted content using user provided schemas
- sudorandom 3mo agoThe buf CLI can do this. Here's how it looks to convert from encoded protobuf to json (protojson). buf convert schema.proto \ --type YourMessageName \ --from payload.binpb \ --to output.json And just invert the arguments to convert back. buf convert schema.proto \ --type YourMessageName \ --from data.json \ --to encoded.binpb
- fisian 3mo agoThere is the --encode (and --decode) options for the `protoc` executable, e.g. `cat myobject.json | protoc --encode MyObject protofile.proto > myobject.bin`
- physicalecon 3mo ago[dead]
- adsharma 3mo agoNeed something for backward compatible RPC? Use protobuf. But for storage and types? This is where I feel there are better alternatives. If you go with protobuf, you're picking 2 out of 3: simplicity, fast, idiomatic. Please consider "uvx tsc-py --help" for things that don't fit.
- adsharma 3mo agoJust in case you find the earlier message too cryptic. It's not performing well on web search. https://pypi.org/project/tsc-py/ https://pypi.org/project/tsc-py/
- jtbaker 3mo agoBuf.build, never again. Had a proto evangelist get us ingratiated into their system a few years ago and then they introduced a bunch of rent seeking behavior just to be able to to the code generation, and getting it ripped out was a huge PITA. https://github.com/betterproto/python-betterproto2 https://github.com/betterproto/python-betterproto2 was a decent (async, type hinted) python implementation last I checked.
- jdpedrie 3mo agoYou can use their CLI and registry for free. I agree that their paid offerings are crazy expensive, but it's simple to just commit your protos in your monorepo (or a proto repo), install buf, create buf.gen.yaml with a few language plugins, and you're up and running. All the pain of protoc is handled. It's been a massive improvement in usability since they came along.
- Scaevolus 3mo agoif you happen to regenerate your protos 5 times in an hour you'll get ratelimited it's a very frustrating system. you should spend a modicum of effort figuring out how to generate your things locally yourself instead, be that docker, arcane wasm...
- krullin 3mo agoinexperienced protobuf user here -- why do non-googlers generally use protobuf? is it for the wire format, or for the schema language, or for the codegen? i have a feeling there is an unmet need for something without all the google-cruft but still gives you a nice schema language and convenient codegen. just a simple stripped-down service and types definition with some support for codegen plugins would cover most crud-like applications in the wild without requiring much fuss. even just supporting only JSON would be acceptable -- maybe an optional buyin for msgpack. i'm thinking twirp, but abandoning protobuf and using a new, simpler schema format.
- seanhunter 3mo agoPeople who want protobuf but better often settle for something like cap’n proto https://capnproto.org/ https://capnproto.org/
- Chu4eeno 3mo agodidn't the capnproto author end up creating protobuf2? or am I mixing things up.
- seanhunter 3mo agoHence the “something like”. There are a few of these sorts of things. I’m not really in the market for them so I don’t try to keep up.
- kentonv 3mo agoThat's me. Other way around -- proto2 came first.
- BobbyTables2 3mo agoCargo-cult mentality. If protobuf hasn’t been created by Google, nobody would use it. It’s actually not well designed for small requests, and doesn’t even deal with versioning all that well. It’s more verbose than a C struct yet not self-describing either. To me, it fails to excel at any one thing yet people blindly use it anyway. If one really considers the details, Protobuf makes ASN.1 actually start to look good!
- parhamn 3mo agoWhats the case for protobuf these days? I loved it for a while. Don't dislike it particularly now (but I've rolled off). - Performance? Fastest JSON lib in most languages are as fast if not faster - Cross language schema generation? So many tools for this these days, they do one thing do them right (depending on your choice of 'right' for things like unions/enums/etc) - The wire protocol? Seems to get in the way vs http2/3. Need special considerations for your proxies (be it nginx or cloudflare). Forced certs, etc are annoying too. I feel like to most shops these days its mostly a schema manager? Protobuf is super bloated for that use case.