8 ms·
Pb-jelly – A Protobuf code generation framework for Rust, developed at Dropbox
- q3k 6y agoFrom pb-jelly-gen: > The core of this crate is a python2 script codegen.py that is provided to the protobuf compiler, protoc as a plugin. That's... surprisingly janky. Not only Python tooling is always painful to deal with (compared to Go/Rust/...), but Python 2? And in a project that otherwise has no reason to depend on Python? :( In comparison, the Go protoc plugin is written in Go, the alternatice rust-protobuf protoc plugin is witten in Rust, the Typescript one is written in Typescript...
- eggsnbacon1 6y agothis might be one of those projects that gets open sourced with the hope that someone else will maintain it
- cbhl 6y agoThe very first prototypes of Dropbox were Py2, so I suspect that legacy is a reason why it was chosen for the codegen. My $dayjob also started as a large python code base, and so there are still lots of one-off Py2 scripts that need to be deleted or rewritten.
- rvz 6y ago> In comparison, the Go protoc plugin is written in Go, the alternatice rust-protobuf protoc plugin is witten in Rust, the Typescript one is written in Typescript... So they're maintaining a project with two languages with Python 2 as a hard requirement for code generation? Oh dear. This tells me that this project will have the same fate as djinni [0] or their similar archived projects. [0] https://github.com/dropbox/djinni https://github.com/dropbox/djinni
- nipunn1313 6y agoHi! One of the authors here. This was an oversight in the documentation. The codegen is py2 and py3 compatible. Fixed! See issues https://github.com/dropbox/pb-jelly/issues/37 https://github.com/dropbox/pb-jelly/issues/37 and https://github.com/dropbox/pb-jelly/issues/40 https://github.com/dropbox/pb-jelly/issues/40 for context.
- rbtying 6y agoanother former contributor to pb-jelly, though no longer at Dropbox. protoc plugins have an interesting bootstrapping problem as well: the protoc-gen-$LANG interface requires the ability to ser/de protobuf messages that describe the proto file's AST. If your build system builds almost everything from scratch, including the protoc plugin, this means that you need to have a variant of your protoc plugin linked to a working proto implementation... That's not to say this is impossible or even difficult, but at the time that I last looked at it (more than a year ago at this point), it made it fairly unpalatable to move the codegen from Python to Rust.
- deepsun 6y agoThere's already 6 different protobuf libraries for Rust: [1] I've chosen Prost for our project, but see the whole list: https://github.com/stepancheg/rust-protobuf#related-projects https://github.com/stepancheg/rust-protobuf#related-projects
- q3k 6y agoIt's 'only' really 4 (two of them are gRPC implementations). But yeah - this is one of those things that makes me stay with Go instead of moving over to Rust for my backend SOA/microservice work. In Rust, for everything you need to do, there's at least 5 different libraries that implement that, all competing with eachother. This is especially annoying when dealing with transitive dependencies. Meanwhile in Go, you generally get one choice - it might be not great, but that's fine, it doesn't have to be. EDIT: This is not intended to be mindless bashing of Rust. I do use Rust for other things. It's a fine language.
- jeffbee 6y agoThere are at least two Go protoc plugins, protobuf-go and gogoproto.
- q3k 6y agoIt's the only real alternative implementation (and more precisely, a fork of upstream protobuf), _and_ there is strong cooperation [1] [2] between both projects to maintain a level of interoperability. And, it's not even protobuf I have a problem with - but things like HTTP implementations. There still isn't a canonical HTTP client/server implementation for Rust, while in Go basically everyone just uses `net/http`, or something that builds on top of that. Same for cryptographic primitives, TLS, context, ... [1] - https://docs.google.com/document/d/19kfhro7-CnBdFqFk7l4_HmwaH2JT_Rhw5-2FLWLEGGk/edit# https://docs.google.com/document/d/19kfhro7-CnBdFqFk7l4_Hmwa... [2] - https://github.com/gogo/protobuf/issues/386 https://github.com/gogo/protobuf/issues/386
- jeffbee 6y ago
- deleted 6y ago[deleted]
- staticassertion 6y agoThis is great. AFAIK this is the only protobuf library in Rust that supports zero copy. Maybe this'll help some of the other libraries implement similar features?
- rapsey 6y agoQuick-protobuf has been around for a while and supports it.
- staticassertion 6y agoThanks, I hadn't seen this. I'm curious about how they compare - it looks like quick-protobuf uses Cow, but I think pb-jelly doesn't?
- rapsey 6y agoJelly uses an external crate called Bytes to achieve it. It has nicer usability compared to Cow I guess.
- lightgreen 6y agorust-protobuf supports zero-copy when using `bytes` feature and reading from `Bytes`.
- staticassertion 6y agoVery cool, TIL.
- jl2718 6y agoIs there really any good reason for code-gen? Just because “google did it” doesn’t mean it’s a good idea.
- q3k 6y agoType safety for your serialization/RPC layer.
- jeffbee 6y agoProtobuf does not provide any type safety whatsoever. The name of the type of the message is carried in a side-channel, and the interpretation of that name is completely up to the endpoint that deserializes the message.
- q3k 6y agoOnce you dispatch your binary/text protobuf into a proto message type however, you do get type safety, and it makes sense to carry over that type safety to implementing languages. Plus, Any [1] is an effort to standardize the ability to carry the proto type (as a global identifier that can be used to retrieve its schema) alongside its serialized format. [1] - https://developers.google.com/protocol-buffers/docs/proto3#any https://developers.google.com/protocol-buffers/docs/proto3#a...
- jakeva 6y agoHow about because for deterministic output, it suits the problem? Just because you can do it by hand doesn't mean I want to.
- foolfoolz 6y agocode generation is great for sharing data models and writing clients for services without dependencies
- SOLAR_FIELDS 6y agoExactly this. Imagine you work at BigCorp and use protobuf to pass data around - now you have a unified data model you can share and everyone can use the same client to access it without going through the trouble of maintaining all those getters and setters. Rolling your own getters and setters is fine in a small project but you really see the advantages of the code gen approach once you are dealing with multiple different teams in an org working with the same complex data model. There are definitely some downsides to the approach though, mostly typical problems you would expect with machine generated code. Namely that it’s verbose and if you have a super complex protobuf data model (hundreds or thousands of fields) and want to ship a fatjar or similar bundling of dependencies you can run into some size issues.
- deleted 6y ago[deleted]
- james412 6y agoI know it's common and perhaps even fashionable, but FWIW language like "We take an opinionated stance" utterly puts me off caring about this package It's a piece of software, it has a design that is either fit for purpose or not. When ego becomes entangled in that design process, it's a strong indicator of the kind of experience one might have trying to get fixes or enhancements merged, or even the kind of attitude you'd find when attempting to report a bug.
- ajkjk 6y agoThat's not what the word 'opinionated' means here. It's not any one person's opinion; it's that the project overall takes a stance on an issue rather than leaving everything open for everyone else to figure out. It provides clarity and direction compared to the more difficult situation where every library is completely general. No ego involved at all.
- jspaetzel 6y agoMore context is helpful here. "We take an opinionated stance that every module should be a crate, as opposed to generating Rust files 1:1 with proto files."
- zxv 6y agoCool. It would be interesting to compare benchmarks the RPC latency under heavy load for pb-jelly compared to other RPC methods. Rust has such great support for performant zero copy serialization and de-serialization in various formats (bincode, message pack, cbor, bson). Seeing this for protobuf feels very encouraging.
- haberman 6y agoI work on the protobuf team at Google, and I'm a big fan of Rust, though I haven't written much actual Rust except a bunch of Project Euler solutions. For protobuf in C++, we've been moving more and more in the direction of using arenas for memory allocation. When you parse a protobuf, it creates a tree of objects that are usually all deleted at the same time. Freeing an arena is much, much cheaper than traversing the tree of objects and calling free() on each one. My dream has been that Rust protobuf could support arenas as well as C++, but use Rust's type system to make it all provably correct at compile time (in C++ the lifetime management is inherently manual and unsafe). For absolute top performance, arenas will always beat trees of unique pointers (which I think corresponds to Rust's Box<> type). I don't know Rust's type/lifetime system well enough to know if this is possible. I was looking recently at arenas in Rust and I noticed that Rust's version of placement new seems to be stalled: "Unfortunately the path forward for placement new in Rust does not look good right now, so I've reverted this crate to work more like a memory heap where stuff can be put, but not constructed in place." https://docs.rs/light_arena/1.0.1/light_arena/ https://docs.rs/light_arena/1.0.1/light_arena/ Does anyone know more about this?
- staticassertion 6y agoAFAIK the current method for hacking in placement new is to use something like this: https://github.com/glandium/boxext https://github.com/glandium/boxext Some more context. There used to be a `box` keyword too. https://github.com/rust-lang/rust/issues/50047 https://github.com/rust-lang/rust/issues/50047
- pcwalton 6y agoPlacement new is just an optimization to avoid the initial memcpy from the stack in cases in which LLVM can't work it out itself. I don't believe that placement new ever enables semantics that aren't possible with plain old move semantics.
- haberman 6y agoAh, in that case it sounds like heterogenous arenas are more or less a solved problem in Rust, even if they aren't necessarily 100% optimal. Probably the more difficult piece then is just how to model arena ownership of a tree of objects that all have links between them. We want to guarantee that links to sub-messages remain valid, which we would expect to be true if they are all in the same arena. But I believe Rust allows moving/swapping objects in and out of the arena?