4 ms·
This is pretty great. Some folks in the thread are suggesting using a binary format, and that's certainly a good idea for some. My business sells software tha
by throwawaygal7 5y ago
This is pretty great.
Some folks in the thread are suggesting using a binary format, and that's certainly a good idea for some.
My business sells software that collects moderately sized (10-100TB/day) , and this data is collected as PROTOBUF. The format is great, and coordinating the backend and client side stuff with protobuf is bliss.
However, we store the data in various relational and document databases. Neither of those use protobuf...
Most importantly, it's business critical to be able to stream this data to data lakes where it can be read by humans. None of the options are going to support protocol buffers, you're either going to have to write a parser (impossible for some) or transform to JSON before ingest (fairly expensive due to some poor choices in how to represent various fields).
It was a sound technical decision to use protoc , and a terrible choice for the business.
I think it would have been much better to use avro, still benefit from schemas but push the work of marshaling JSON back to clients...
- wodenokoto 5y agoI know it’s utopia, but I still can’t help but ask, why can’t we have protobuffs or some other binary format with readers integrated into normal systems the way everything has a Json reader. I mean, stuff like “double click on .protobuff file and it opens in notepad and looks like json or whatever. When you click save it gets serialized back to protobuff.
- MayeulC 5y agoIIRC (I could be wrong on this one) protobuf isn't really a specification, it's a single implementation, and is quite hard to reimplement. Here's a short comparison of serializing formats: https://drewdevault.com/2020/06/21/BARE-message-encoding.html https://drewdevault.com/2020/06/21/BARE-message-encoding.htm...
- morelisp 5y agoRegarding the wire format specifically, it's not fully formalized but there's dozen implementations, not all of them from Google. You could hack together a basic serde for a particular language in an afternoon, in some respects much more easily than you could JSON. Most of the engineering work is in the schema compilers.
- heavenlyblue 5y agoBut in order to semantically read protocol buffer files one needs to compile the schema files…
- morelisp 5y agoThe degree you "need" to support the full schema format depends entirely on the language you're using. In Go you only need the annotated structures, or in Java you only need a mapping between field number and name, and the languages' own reflection capabilities can handle the rest. Yes, strings, bytes, and substructures all appear in the same in the wire format. Just like strings, bytes, and dates all appear the same in JSON. If you're trying to write a generic protobuf viewer like wodenokoto suggests, you can make reasonable assumptions about the contents 99% of the time based on the data, and show the user multiple options if you're not sure. There are already lots of tools that do this, but none very well integrated into mainstream development workflows.
- matja 5y agoTricky part about implementing that is that the protobuf wire-format alone doesn't contain enough information to unambiguously represent the real data types, it also needs the schema (.proto file) to do that. For example, in the wire-format, a string and a sub-message are encoded as the same type (a blob - varint + sequence of bytes), but using the schema they are clearly interpreted differently. Sure, it is possible to make a self-describing protobuf message which includes the schema in protobuf representation, but that is a special-case.
- uDontKnowMe 5y agoWhat about Avro then?
- Too 5y agoJson schema has a similar challenge, there is no header, in the json files themselves, indicating which schema is used. The slightly dirty solution is https://www.schemastore.org/json/ https://www.schemastore.org/json/, where the IDE looks up schema from a global registry, using a fileMatch pattern.
- gravypod 5y agoI've written translation layers for such systems and it's not too bad. See this project from $job - 1: https://github.com/CaperAi/pronto https://github.com/CaperAi/pronto It allowed us to have a single model for storage in the DB, for sending between services, and syncing to edge devices.
- throwawaygal7 5y agoThis is great work! I've done similar stuff (not as nice) for some relational and document stores. But if you're talking about data lakes, like splunk, or various other similar systems - there's no standard way for writing a parser like this across all of them and you end up implementing the same thi ng a ton of different times.
- axiosgunnar 5y ago> moderately sized (10-100TB/day) weird flex but ok :-)