3 ms·
I agree with a lot of this post, although the tone isn’t great. The problems we ran into with protobufs at my job include: 1. The schema evolution claims don’t
by ves 7y ago
I agree with a lot of this post, although the tone isn’t great. The problems we ran into with protobufs at my job include:
1. The schema evolution claims don’t really hold water for our systems.
2. The type system isn’t very expressive (e.g. no generics means you have to write the same error wrapper for all your endpoints) and lots of our devs found it unintuitive, especially oneofs.
3. The “default value”/nullable field feature turns out to be a recipe for postmortems and data quality degradation. Making everything nullable isn’t good.
4. The python library doesn’t have mypy typing and the generated objects aren’t... super pythonic.
I (along with some colleagues) built a library to paper over protobuf and address these issues. Notably, it includes a very well-specified algorithm to automatically assign version numbers to schemas during development, as well as provide operational instructions to avoid bumping a version without causing downtime if possible. And all the codegenned models have mypy types!
You can read more about it here:
https://tech.affirm.com/defining-data-models-with-idol-a31093cd0707 https://tech.affirm.com/defining-data-models-with-idol-a3109...
It’s so far turned out really really well for us.
In particular, “schema evolution” is a property of a particular distributed system and there aren’t universally safe rules; schemas for historical machine learning datasets and rpc services, say, have to evolve differently cos the data flow is different. Also, there’s no version bumping algorithm built in, and nullable/optional fields are a pain to program against for data scientists and client devs alike.
- GeneralMayhem 7y agore: (3) - nullability is more or less required for backwards compatibility. If you have existing data and add a new field going forward, your options are to make the old data invalid until you backfill, or give your code a way to detect "this field doesn't exist" and deal with it accordingly.
- deleted 7y ago[deleted]
- ves 7y agoI opted to go for “pinning” based on the version number, so if you make a breaking change, like adding a required field, IDOL copies your schema into a v2 (say) namespace and then applies the change, leaving v1 untouched. At this point we just have separate types for separate versions and tools in the host language can help you deal with that. This turns out to be much better for data quality and client code than adding lots of nullable fields, at the cost of making breaking changes to APIs a bit more work. It seems to have been worth it so far. Going forward, the service author has to support the “old” versions until we can determine that there’s no old data sitting around (so all clients are on the new version, all serialized data has been backfilled or dropped, or whatever’s appropriate), at which point they can delete the old schema. And we have some simple tools to verify this, since we stick the version number onto the models / serialized data.
- kortex 7y agoYou can generate stubs for mypy with this tool but yes this should be something supported out of the box. https://github.com/dropbox/mypy-protobuf https://github.com/dropbox/mypy-protobuf
- vivekseth 7y agoCouldn’t find IDOL on Affirm’s github. Is any of it open source? I’d love to take a look if so.
- ves 7y agoIt isn’t... yet! it would be a bit of work to open source it (rip out any lingering affirm bits, spruce it up some) and my team is real pressed for time at the moment. but I definitely want to do it sooner rather than later!