5 ms·
If by "solving" you mean "refuse to do anything at all unless you have the exact schema version of the message you're trying to read" then yes. In a RPC context
by dtech 3y ago
If by "solving" you mean "refuse to do anything at all unless you have the exact schema version of the message you're trying to read" then yes. In a RPC context that might even be fine, but in a message queue...
I will never use Avro again on a MQ. I also found the schema resolution mechanism anemic.
Avro was (is?) popular on Kafka, but it is such a bad fit that Confluent created a whole additional piece of infra called Schema Registry [1] to make it work. For Protobuf and JSON schema, it's 90% useless and sometimes actively harmful.
I think you can also embed the schema in an Avro message to solve this, but then you add a massive amount of overhead if you send individual messages.
[1] https://docs.confluent.io/platform/current/schema-registry/index.html https://docs.confluent.io/platform/current/schema-registry/i...
- insanitybit 3y ago> but it is such a bad fit that Confluent created a whole additional piece of infra called Schema Registry [1] to make it work. That seems like a weird way to describe it. It is assumed that a schema registry would be present for something like Avro. It's just how it's designed - the assumption with Avro is that you can share your schemas. If you can't abide by that don't use it.
- dtech 3y agoI do not think its unfair at all. Schema registry needs to add a wrapper and UUID to an Avro payload for it to work, so at the very least Avro as-is is unsuitable for a MQ like Kafka since you cannot use it efficiently without some out-of-band communication channel.
- insanitybit 3y agoEveryone knows you need an out of band channel for it, I don't know why you're putting this out there like it's a fault instead of how it's designed. Whether it's RPC where you can deploy your services or a schema registry, that is literally just how it works. Wrapping a message with its schema version so that you can look up that version is a really sensible way to go. A uuid is way more than what's needed since they could have just used a serial integer but whatever, that's on Kafka for building it that way, not Avro.
- morelisp 3y ago> a serial integer And now you can't trivially port your data between environments.
- insanitybit 3y agoCan you elaborate? I don't see any issue at all.
- morelisp 3y agoUnderstanding the data now depends not just on having a schema to be found in a registry, but your schema registry, with the schemata registered in the same specific order you registered them in. If you want to port some data from prod back to staging, you need to rewrite the IDs. If you merge with some other company using serial IDs and want to share data, you need to rewrite the IDs. Etc.
- insanitybit 3y agoI have no idea what you're talking about. If the standard were a 64bit integer none of what you're saying would be the case at all. There is no difference between a random 128bit integer vs a sequentual 64bit integer except that the 64bit integer is smaller.
- kentonv 3y agoThey're saying that for a sequential number, you must have a central authority managing that number. If two environments have different authorities, they have a different mapping of these sequential numbers, so now they can't share data.
- insanitybit 3y agoOh. Like, if you have two different schema registries that you simultaneously deploy new schemas to while also wanting to synchronize the schemas across them. Sounds weird as hell.
- nly 3y agoHaving the schema for a data format I'm decoding has never been a problem in my line of work, and i've dealt with dozens of data formats. Evolution, versioning and deprecating fields on the other hand is always a pain in the butt.
- dtech 3y agoIf a n+1 version producer sends a message to the message queue with a new optional field, how do the n version consumers have the right schema without relying on some external store? In Protobuf or JSON this is not a problem at all, the new field is ignored. With Avro you cannot read the message.
- nly 3y agoI mean a schema registry solves this problem, and you just put the schema in to the registry before the software is released. A simpler option is to just publish the schema in to the queue periodically. Say every 30 seconds, and then receivers can cache schemas for message types they are interested in.