12 ms·
Cap'n Proto 0.6 released – 2.5 years of improvements
- hellofunk 9y agoThere is some strange geeky enjoyment from browsing all the serialization libraries out there. For my taste, I have centered on Cereal. When workind with C++ end to end, I have found it to be the easiest and fastest way to throw data around.
- Heliosmaster 9y agoSerialization can easily be a bottleneck, especially for write-heavy systems (like we do, with Event Sourcing). So i think it's quite natural that people try and squeeze the last few milliseconds out of it. It can easily make a huge difference when serializing/deserializing billions of entities
- fh973 9y agoSerialization is order of microseconds, not milliseconds.
- spullara 9y agoMicroservice architecture CPU usage can be dominated by serialization. It is one of the reasons that the JVM was so much faster than Ruby at Twitter for the frontend. The business logic just didn't matter as much as deserializing thrift and serializing json or html.
- pagnol 9y agoThere's a Wikipedia article that provides a nice overview: https://en.wikipedia.org/wiki/Comparison_of_data_serialization_formats https://en.wikipedia.org/wiki/Comparison_of_data_serializati...
- educar 9y agoWindows support is great but why is this a showstopper? This is a common problem in many libraries and the common solution is to mark those platforms as unsupported until sometime steps up. Also platform support is a dynamic list - platforms are kept alive by presence of contributors/maintainers.
- kentonv 9y agoI think it would set a pretty bad precedent to have Windows support in one release and then drop it in the next release, then bring it back, etc. Windows users would likely be left very confused. Saying "it's up to contributors to step up" is nice in theory but in practice I don't think a well-maintained library can operate that way. Few people are eager to volunteer to run tests over and over fixing tiny issues for a coordinated release... That said, it probably would have been better to drop Windows temporarily than to go two years without a release. But it always seemed like I'd find time in the next month. Also note that Windows wasn't the only thing. I have a huge test matrix that I run for every release and, again, to avoid confusion, I want the whole thing to pass for any release. Things like building with -fno-exceptions or 32-bit builds or Android or ancient GCC versions tend to break frequently as the code evolves, but we really ought to have all these working for a release. But maybe I'm just too OCD about this... :) We now have AppVeyor (and Travis-CI) set up to build a chunk of the test matrix on every commit, which should help a lot going forward.
- speps 9y ago> But maybe I'm just too OCD about this... :) It's not OCD, it's good practice and I wish more open-source projects would follow that! Good job!
- justin66 9y agoKenton, the Windows support looks amazing and I'm grateful for it. There are so many of us for whom lack of up to date Windows support is a showstopper, so, thanks! (Thanks to Harris and others, as well)
- ipsum2 9y agoGreat stuff! Since this release comes after the release of GRPC and (slightly less related) Graph API and ships with a web server, how does it compare?
- kentonv 9y agoCap'n Proto RPC was originally released before gRPC. That said, gRPC is heavily based on Google's internal RPC system which has been around for a very long time, albeit not publicly. If you read the Cap'n Proto RPC docs, everywhere where it mentions "traditional RPC", I specifically had Google's internal RPC in mind (having previously been the maintainer of Protobufs at Google). So, you can more-or-less substitute gRPC in there for a direct comparison. https://capnproto.org/rpc.html https://capnproto.org/rpc.html There are two key differences: 1. Cap'n Proto treats references to RPC endpoints as a first-class type. So, you can introduce a new endpoint dynamically, and you can send someone a message containing a reference to that endpoint. Only the recipient of the message will be able to access the new endpoint, and when that recipient drops their reference or disconnects, you'll get notified so that you can clean it up. This is incredibly useful for modeling stateful interactions, where a client opens an object, performs a series of operations on it, then finally commits it. Put another way, this allows object-oriented programming over RPC. Also note that you can easily pass off object references from machine to machine -- currently this will set up transparent proxying, but in the future we plan to optimize it so that machines automatically form direct connections as needed, which will be really powerful for distributed computing scenarios. 2. Relatedly, Cap'n Proto supports "promise pipelining", which allows you to use the result of one RPC as an input to the next without waiting for a round-trip to the client. This makes it possible to use object-oriented interaction patterns with deep call sequences without introducing excessive round-trip latency. This is described in detail at the RPC link above.
- u320 9y ago> Put another way, this allows object-oriented programming over RPC. So... CORBA?
- 9y ago
- forrestthewoods 9y agoWoohoo! Lack of first class Windows support always held me back. Looking forward to playing with this and seeing hopefully more regular future updates.
- pagnol 9y agoSome time ago I wanted to use Cap'n Proto in the browser but then I found that the only existing implementation written in JavaScript hadn't been updated in two years and the author himself recommended against its use somewhere in a thread on Stackoverflow. I would love to use Cap'n Proto but for me a robust JS implementation is a sine qua non. Does anyone here happen to know if there's been any progress in this regard or have I missed something?
- kentonv 9y agoThere has been some work on a new Javascript implementation, as described in this thread: https://groups.google.com/d/msg/capnproto/lESKRE_pix8/jX59zEUXBgAJ https://groups.google.com/d/msg/capnproto/lESKRE_pix8/jX59zE... But it's been slow. On the other hand, now that 0.6 includes a JSON library, it's relatively easy to do browser<->server in JSON and then use Cap'n Proto on the back-end. But, obviously, it would be nicer to use Cap'n Proto through the whole stack. Contributions are welcome!
- kybernetikos 9y agoI've got about 70% of a capnproto implementation written in pure javascript that works in the browser. Part of the problem is that there is a bit of an ideological difference, where I prefer dynamic code to adding build steps but all the existing tools and code assume you want to generate source. I also found it quite irritating to bootstrap too, because the tooling itself uses capnproto. It sounds neat, but then it means that you have to be able to read capnproto in order to be able to read it. The documentation was also incredibly patchy at the time which made it quite hard to develop for, although I'm told that the kentonv & gang are very helpful. In the end, I was just starting to struggle with capnproto generics when my requirement went away as the server I was connecting to added a msgpack option.
- pagnol 9y agoDo you happen to have a repository for this on Github?
- 9y ago
- grandinj 9y agoDoes it do protocol negotiation? i.e. can a client ask the server what interfaces it implements?
- kentonv 9y agoThere are a few answers to that: 1. If you have a remote object reference, you can (explicitly) cast it to any interface type, and then attempt to call it. If it doesn't implement that interface (or that method), an "unimplemented" exception will be thrown back. It's relatively common to do feature detection this way. 2. It's easy for the application to define a Cap'n Proto interface to support fancier negotiation. For example, you could have all your RPC interfaces extend a common base interface which has a method getSchema() which returns the full interface schema, or a list of interface IDs, or whatever it is you want. 3. We actually plan to bake in schema queries in a future version, such that all objects will support some sort of getSchema() call implicitly. This would especially be useful for clients in dynamic languages that could potentially connect to a server without having a schema at all, and load everything dynamically.
- drej 9y agoThis might be a contrarian view, maybe I'm misunderstanding it. To me, much of message passing is not performance critical, so it would be well served by JSON/YAML/XML for easy implementation/testing/debugging. If one needs performance, he can just send bytes over, which can be (de)serialised in a (couple) dozen SLOC. Sure, when you're talking about very complex structures, RPC, dynamic schemas etc., then you might opt for something like this, but let's be honest - that's quite a minority of current users, isn't it? I never minded these frameworks, but then I wanted to write a few parsers for some file formats and they used Thrift/Flatbuffers to encode a few ints, which seemed like a major overkill. There was no RPC, no packet loss, no nothing.
- willvarfar 9y agoClearly there are needs for very efficient message parsing. That's why these serialization frameworks exist. I've got servers that spend too many cycles in serialization, so I pay careful attention to it. I put a lot of effort into using JSON despite the performance problems, and have pretty low-level code to read write it without too many allocations etc. Nothing thats easy on the eyes. But I'd hazard a guess that you're spot on regards the normal needs of normal programs. The vast majority of programs utilizing message-passing are not serialization-bound. For these programs I'd recommend JSON because it is easy for humans to inspect and debug, and there are plenty of libraries to choose from if you are not performance-sensitive. (YAML is hopeless regards security and XML is just horrid to look at when you're debugging).
- bborud 9y agoBeing able to inspect messages is down to the proper tooling. But I can certainly understand where you are coming from. I've designed several text based protocols in my time for that exact reason. And I still do for stuff that I positively know will never require any sort of performance (mostly on embedded systems, somewhat ironically since those are sometimes constrained down to just a few hundred bytes of memory :-))
- edem 9y agoWhen you start having polyglot microservices this all will make sense. For example protobuf gives you versioning and compiles to a dozen languages...just because it is not useful to you don't dismiss the idea. And since I'm part of the minority I'm happy to have libs like this around. Just a sidenote: this behavior of neglecting performance alltogether leads to systems which take ages to build and collapse under pressure.
- TeeWEE 9y agoVer nice and technically better than gRPC. However i better bet my company success on a standard that the big giant Google is betting on, and other companies are embracing. Also client side libraries for all langauges is important. Probably gRPC is probably better here too?
- ocdtrekkie 9y agoFWIW, Cloudflare is betting on Cap'n Proto, and Cloudflare is pretty big too.
- kentonv 9y agoThough to be fair, Cloudflare is still 2-3 orders of magnitude smaller than Google by most measures. :)
- DonHopkins 9y agoThe name "Cap'n" was forever tainted for me, from my traumatic experience with "Cap'n Software Forth". http://www.art.net/~hopkins/Don/lang/forth.html http://www.art.net/~hopkins/Don/lang/forth.html "The first Forth system I used was Cap'n Software Forth, on the Apple ][, by John Draper. The first time I met John Draper was when Mike Grant brought him over to my house, because Mike's mother was fed up with Draper, and didn't want him staying over any longer. So Mike brought him over to stay at my house, instead. He had been attending some science fiction convention, was about to go to the Galopagos Islands, always insisted on doing back exercises with everyone, got very rude in an elevator when someone lit up a cigarette, and bragged he could smoke Mike's brother Greg under the table. In case you're ever at a party, and you have some pot that he wants to smoke and you just can't get rid of him, try filling up a bowl with some tobacco and offering it to him. It's a good idea to keep some "emergency tobacco" on your person at all times whenever attending raves in the bay area. My mom got fed up too, and ended up driving him all the way to the airport to get rid of him. On the way, he offered to sell us his extra can of peanuts, but my mom suggested that he might get hungry later, and that he had better hold onto them. What tact!"
- mhogomchungu 9y agoI think this[1] API among others should be renamed and "future" be used instead of "promise" to better reflect C++ naming conventions. [1] https://github.com/sandstorm-io/capnproto/blob/247e7f568b1664cbcca1455243da06a7ab78c6a3/c%2B%2B/src/kj/async-io.h#L54 https://github.com/sandstorm-io/capnproto/blob/247e7f568b166...
- kentonv 9y agoCap'n Proto Promises are highly analogous to promises in Javascript. Both are derived directly from the E language, which has been around for decades. It makes sense to stick with the terminology used in similar designs, so that people don't have to re-learn the same concept when switching languages. C++'s standard library for some reason decided -- relatively recently -- to introduce the term "promise" to mean something different (what most people call a "resolver" or a "fulfiller" for a promise), which is unfortunate. It's C++ that is being inconsistent here. There are also subtle historical differences between the meaning of "promise" and "future". Historically, futures have usually existed in multi-threaded designs rather than event-loop/callback designs; you wait() on a future, blocking the calling thread, whereas with promises you call promise.then() to register a callback to call when a promise completes. The C++ committee again seems to have gotten this confused with their definition of "future", which looks more like a traditional promise.
- sluukkonen 9y agoI have no idea why the C++ designers chose those particular names, but at least Scala uses the same Promise/Future dichotomy, where you resolve (or reject) a Promise into a Future. Perhaps they drew from the same well?
- moosingin3space 9y agoHow do Cap'n Proto Promises and Rust futures relate?
- dwrensha 9y agoThey have a lot in common! For a while, capnp-rpc-rust used `gj::Promise`, which is based directly on the C++ Cap'n Proto implementation of promises (i.e. `kj::Promise`). Back in January, capnp-rpc-rust was updated to use `futures::Future` instead, and it was a fairly straightforward transition, as described in this blog post: https://dwrensha.github.io/capnproto-rust/2017/01/04/rpc-futures.html https://dwrensha.github.io/capnproto-rust/2017/01/04/rpc-fut... The trickiest part of the transition was dealing with scheduling. The implementation of `kj::Promise` has a built-in scheduling queue that guarantees a certain form of deterministic FIFO semantics, and those semantics are heavily depended upon in the Cap'n Proto RPC implementation. Rust's `future::Future` is less batteries-included, requiring capnp-rpc-rust to explicitly create queues where deterministic scheduling is needed. Confusing the terminology perhaps even more, in capnproto-rust there is a type `capnp::capability::Promise` that implements `futures::Future`.
- halestock 9y agoA bit off topic, but that title banner is a massive waste of space. It takes up more than half the space on my screen (15" laptop).
- taftster 9y agoI know, totally agree. It's actually one of those things that have made me shy away from recommending its use. I continually end up back at Protocol Buffers, simply because the Cap'n Proto website doesn't seem professional to me. Instead, it feels like the back of a breakfast cereal box.
- DonbunEf7 9y agoAs usual, nobody has said "capability" yet, which is unfortunate, because one of Capn's biggest strengths is that it embodies the object-capability model and is cap-safe as a result. Edit: Why does this matter? Well, first, it matters because so little software is capability-safe. Capn's RPC subsystem is based directly upon E's CapTP protocol. (E is the classic capability-safe language.) As a result, the security guarantees afforded by cap-safe construction are extended across the wire to cover the entire distributed system. This security guarantee holds even if not every component of individual nodes is cap-safe, and that's how Sandstorm works. Continuing on, there's also historical stuff going on. HN loves JSON; JSON is based on E's DataL mini-language for data serialization. HN loves ECMAScript; ES's technical committee is steered by ex-E language designers who have been porting features from E into ES.
- placeybordeaux 9y agoCould you elaborate?
- pas 9y agoI think the built in security of "you can't call if you don't know the address" coupled with addresses being first-class (so you can send the address of a function to a client, hereby granting the capability): https://news.ycombinator.com/item?id=14244540 https://news.ycombinator.com/item?id=14244540 this is very similar to what the new bus1 ipc author wants to do: https://www.youtube.com/watch?v=6zN0b6BfgLY https://www.youtube.com/watch?v=6zN0b6BfgLY
- kentonv 9y agoThe trouble is, we capability people have come up with our own language full of jargon that no one else knows, and I think it confuses people. When talking to people new to the idea, I try to use the term "object reference" or maybe "endpoint reference" rather than "capability", to be more approachable.
- ludwigvan 9y agoIs this the E language you are referring to? Had never heard of it before. https://en.wikipedia.org/wiki/E_(programming_language) https://en.wikipedia.org/wiki/E_(programming_language)
- makmanalp 9y agoWhat always trips me up about capnproto is that it's billed as a serialization library, but what it is is an in-memory storage layout, and "serialization" is mostly just dumping memory into a file, right? (which is cool) What confuses me is, then what are the costs of migrating to this system? Am I essentially dumping my programming language's object model for my capnproto implementation's? When can this be annoying? Or does it vary from implementation to implementation? In a similar tangent - how similar is this to apache arrow, not because of the columnar analytics part, but could I expect to just dump a bunch of data in shared memory and read it from another process to eliminate IPC serialization/copy costs?
- kentonv 9y agoGenerally I'd recommend using Cap'n Proto in much the same way as you'd use Protobuf. It's not intended that your in-memory state be in Cap'n Proto objects, only the messages you intend to transmit, or data stored on disk.
- makmanalp 9y agoWait, but then aren't you merely shifting the serialization cost instead into building capnproto objects? (and perhaps that's more efficient somehow?) It seemed to make more sense when you already have your data as capnproto objects, versus creating objects only to send and discard them, which is similar to regular old serialization again.
- kentonv 9y agoWith Protobuf, you still have the cost of constructing the objects, and the cost of then serializing them. Cap'n Proto removes the latter cost. Serializing is generally the much more expensive step. It also turns out Cap'n Proto reduces the building cost by a fair amount, because the arena-style allocation needed to support zero-copy output also happens to be a lot cheaper and more cache-friendly, but that's somewhat of an accident. For message-passing scenarios, Cap'n Proto is an incremental improvement over Protobufs -- faster, but still O(n), since you have to build the messages. For loading large data files from disk, though, Cap'n Proto is a paradigm shift, allowing O(1) random access.
- justforFranz 9y agoWould it kill developers of open source to provide a simple, two-sentence explanation of what their app does?
- paulddraper 9y agoHow's this? > Cap’n Proto is an insanely fast data interchange format and capability-based RPC system. Think JSON, except binary. Or think Protocol Buffers, except faster. https://capnproto.org/ https://capnproto.org/
- kentonv 9y agoI guess I should make the top banner be a link to the home page, so that people on mobile (who can't see the sidebar very easily) can just mash the screen. EDIT: done