9 ms·
Prototool – A Swiss Army Knife for Protocol Buffers
- grizzles 8y agodanby - grpc for the browser :: is looking for testers https://github.com/ericbets/danby https://github.com/ericbets/danby There are two upcoming features. The first one is streaming support. The second is a callback API template that mirrors the grpc node API exactly. Or you will have the choice to stick with the current promise API. It's not a priority for us at the moment but adding a simple load balancer that distributed traffic randomly across a set of servers would be a ~5 line patch.
- throwaway84742 8y agoIn another decade or so the world might replicate half of the very nice internal tools Google has. Suggestion for a project: make a tool that, given a proto description and a file that contains concatenated proto messages stored as binary strings (sort of like RecordIO at Google) lets you run simple SQL queries on the data and extract a subset of the fields from messages matching a predicate, and maybe even do simple aggregations. That was pretty handy. I really wish Google would open source some or most of this stuff. It’s not like keeping it closed source creates any kind of insurmountable competitive advantage, especially compared to the advantages that would accrue from broader adoption of protobufs.
- Willson50 8y agoYou might be interested in KSQL, SQL queries that run on Kafka streams. https://www.confluent.io/product/ksql/ https://www.confluent.io/product/ksql/
- throwaway84742 8y agoNah. I’m interested in quickly querying on-disk data specifically, ie proto-based application logs and the like (another thing the world needs to adopt more broadly imo).
- pedge 8y agoOf note, Prototool has a binary-to-json command, so assuming your Protobuf files are in path/to/proto/files, and you had a newline-separated log file of Protobuf messages foo.bar.Baz, you could do: cat /path/to/log.file | prototool binary-to-json path/to/proto/files foo.bar.Baz - | jq .search.term
- puzzle 8y agoOther tools and features that don't exist outside: - a tee loadbalancer for gRPC, forwarding the same requests to both A and B backend pools, but only returning results from A. I don't think Envoy has this, but it should. - load balancing dashboards showing traffic between frontends and backends - load balancer support for dynamic sharding - gnubbyd under ChromeOS: https://groups.google.com/a/chromium.org/forum/m/#!msg/chromium-hterm/LuDVJ67Q4BE/vKoY0j_RBKQJ https://groups.google.com/a/chromium.org/forum/m/#!msg/chrom... (I think most of this is doable these days, but the initial setup requires a Linux system) - Kubernetes: server-specific custom hyperlinks on dashboards (e.g. links to POD_IP:PORT/stats, /debug, etc. for each individual pod you are looking at) - Kubernetes: multiple Docker images in the same container or pod. E.g. the first container could be your code, while the second one might be data or the JVM runtime, etc., without having to bundle them together or doing costly copies in init containers. - Kubernetes: canaries and automatic rollbacks
- kodablah 8y ago> Kubernetes: canaries and automatic rollbacks Hot off the presses: https://cloudplatform.googleblog.com/2018/04/introducing-Kayenta-an-open-automated-canary-analysis-tool-from-Google-and-Netflix.html https://cloudplatform.googleblog.com/2018/04/introducing-Kay.... Though you have to use Spinnaker.
- puzzle 8y agoThat's an external controller, which is what most people are doing themselves these days, reinventing the wheel each time. Borg has long had an automatic rollback feature on updates, tuned through a few settings on top of the health check machinery. I'm in the camp believing that a basic implementation should be built-in, since health checks are already there. An implementation of this was started, but it has stalled. External controllers should be still allowed, for more advanced scenarios like rolling back upon detection of regressions in latency and similar. Also, setting up Spinnaker is pretty much as complicated as Kubernetes itself. :-)
- lobster_johnson 8y ago
- zellyn 8y agoWhen I was at Google, I kept an eye on the open sourcing of RecordIO. Apparently there was no desire not to open source it: it was simply that nobody had the time to disentangle and/or clean it up for release. Looks like some parts of it have escaped… https://github.com/eclesh/recordio https://github.com/eclesh/recordio
- vinkelhake 8y agoIf you were interested in RecordIO, then this project might also be of interest to you: https://github.com/google/riegeli https://github.com/google/riegeli
- zellyn 8y agoInteresting. I wonder how different that is from RecordIO. Also, whether there'll be a Go implementation. [Edit, after looking a bit.] Pretty different. If I remember correctly, RecordIO is re-synchronizing, whereas Riegeli seems to break things up into 64KB chunks, splitting messages across chunks if necessary. [Edit, after finding more information.] Interesting… looks like Riegeli is intended to compress well, rather than just store sequentially. https://encode.ru/threads/2895-Riegeli-%E2%80%94-a-new-compressor-for-structured-data-(protocol-buffers) https://encode.ru/threads/2895-Riegeli-%E2%80%94-a-new-compr...
- throwaway84742 8y agoIIRC (but memory has faded considerably), RecordIO also did support something to aid compression across records (rather than just offer per-record compression). There was some gnarly code in it to that effect where there could be a compressed subset of several records within the file. But I might be wrong.
- throwaway84742 8y agoPretty neat work. Especially the corruption detection/skipping and seek support. That part was super ugly in RecordIO proper, and relied on stars properly aligning and the absence of cosmic radiation. RecordIO being the default format for everything at Google, it's not something they can really fix though. Taking care of concatenation is a nice touch as well. As someone who has spent quite a bit of time at Google working on a high performance file format (not RecordIO): 1. I'd also add LZ4 and/or Snappy for the cases where they are more Pareto-optimal (i.e. fast, network attached, remote storage, such as SSD Colossus, or its external proxy: SSD Persistent Disk). 2. IMO HighwayHash is overkill here, and the author should have used CRC32C instead. You don't particularly care about collisions in this case, you're detecting data corruption. CRC32C is perfect for that, and it's hardware accelerated in almost all recent Intel and ARM CPUs, and it's half the size on disk. 3. It'd be pretty cool to introduce some kind of metadata which would tell the user what type of message is encoded in the file. This is not something RecordIO has, but internal tools can guess most of the time because they have all the proto definitions at their disposal. There's no need to store it in every header, just the first one. I would advise against storing the full schema (that can get very gnarly in the presence of proto dependencies and extensions), but just have something lightweight, i.e. message name and perhaps SCM revision number or hash in the file header, so that the user (or the external system consuming the files) could somewhat reliably establish what the format is later on, when the proto definition drifts. Otherwise, this being a binary serialized file format, it's very easy to end up in a situation where you have some files from years ago and you no longer know how to read them. And yes, I'm aware that SCM hash can change if history is edited.
- grandinj 8y agoYou could add a TableEngine extension to H2 (h2database.com), pretty easily which would give you full SQL query functionality over such a file
- throwaway84742 8y agoNope. Protos have repeated fields and can be hierarchical (that is, can contain other protos) and even recursive (that is, contain themselves, possibly as repeated fields). H2 is not going to work.
- grandinj 8y agoYeah, a normal SQL model is not perfect because of the requirement that it look like a flat table i.e. fixed number of columns. But H2 has ARRAY for repeated fields, and with some custom functions for decomposing other functions, you could get pretty far. Just saying, not perfect, but could be useful without too much effort.
- chrissnell 8y agoI would also love to see a protobuf/gRPC decoder for wireshark. Bonus: the ability to filter sniffed packets based on a field value.
- SOLAR_FIELDS 8y agoWhile not exactly what you are describing, I work for another company that uses protobufs extensively and we have some nice internal tools similar to what you describe. I really wish we could open source those too. I feel like the wheel is reinvented a lot with protobuf in several of the large companies who use it.
- lobster_johnson 8y agoHow does RecordIO compare with Parquet and Arrow? Different use cases?
- throwaway84742 8y agoDon’t know about Arrow, but Parquet is a columnar format. Such formats can’t write record-by-record, they need a large number of records to shred into columns in order to realize their columnar benefits. In contrast, appending to RecordIO is little more than writing a binary string. The downside of RecordIO is that you can’t just read some fields in a message and not others. You have to deserialize the whole message. RecordIO is cheap to write and well suited for cases where reading the entire message is not that big a deal. Columnar formats are more suited for the cases where it’s ok to pay the relatively substantial up front encoding cost for vastly greater performance in analytical workloads. Advanced ones contain additional metadata (such as range and hash constraints, the former can be both per file and per block) which the analytical runtime will be able to take advantage of in order to avoid doing the work that doesn’t need to be done.
- reacharavindh 8y agoHave limited protobuf knowledge. Why not use SQLite[1] for storing this data? Storing structured data in binary format, and being able to run SQL queries on it, is already possible with SQLite right? [1] - https://www.sqlite.org/appfileformat.html https://www.sqlite.org/appfileformat.html
- deleted 8y ago[deleted]
- endymi0n 8y agoI‘m smelling the SQL case could be reasonably easily thrown together with PostgreSQL and a custom Foreign Data Wrapper based on protobuf-c (prior art: cstore_fdw by the Citus folks). Proto definitions then should compile rather cleanly to table definitions, at least one level down (PG isn‘t so good with nested structures). The main thing stopping this endeavour is probably that to the best of my knowledge, there isn‘t any standardization in the Protobuf community about file formats serializing multiple of these together like RecordIO - that, and my C skills are pretty rusty by now :)
- erik_seaberg 8y agoSounds a lot like a Hive query over self-describing Avro files.
- kodablah 8y agoSays "Handle installation of protoc [...] behind the scenes in a platform-independent manner without any work on the part of the user", doesn't support Windows yet [0]. Granted, as pre-1.0 I should probably read the features as goals. 0 - https://github.com/uber/prototool/issues/9 https://github.com/uber/prototool/issues/9
- mabynogy 8y agoI can't use something with a CoC.
- n42 8y agoWhy not?
- recursive 8y agoI'm guessing the code disallows their conduct.
- mabynogy 8y agoIt's a political tool. It means they are into politics.
- zrail 8y agoEverything involves politics. At least they're being explicit about it.
- mabynogy 8y agoNo: physics, maths, cooking...
- zrail 8y agoOh please. Things involving more than one human invariably involve politics. Academics are just as susceptible as anything else, probably more so. Cooking, if there is more than one person involved in preparing and eating the food, will involve politics.
- mabynogy 8y agoNo. You're mixing social interactions and politics. Politics is about power (who is the boss).
- adam_gyroscope 8y agoWe built https://github.com/GyroscopeHQ/grpcat https://github.com/GyroscopeHQ/grpcat at my company, which takes text-format protos as input and sends them to a gRPC endpoint. Looking at Prototool I think I should just merge the functionality into Prototool. This is cool!
- hurricaneSlider 8y agoAlways wished there was a tool for protobuf which could test whether a changes to any .proto files were backwards compatible and if not raise an error
- flippmoke 8y agoHere are some other great tools that is quite useful with protobufs, one in C++ and one in pure javascript. https://github.com/mapbox/protozero https://github.com/mapbox/protozero https://github.com/mapbox/pbf https://github.com/mapbox/pbf
- ris 8y agoYet another tool trying to "manage" "packages" on my machine!
- deleted 8y ago[deleted]