13 ms·
Cap’n Proto
- Twonneilb22ll 10y agoI'm happy to hear from you all in glad to see you're doing all you can for updates Programs codes appreciate the hard work
- joshuawarner32 10y agoHere's the discussion from a while ago: https://news.ycombinator.com/item?id=5482081 https://news.ycombinator.com/item?id=5482081
- deleted 10y ago[deleted]
- throwaway13337 10y agoTo get an overview of the area of binary interchange formats that are language agnostic, the author of Cap'n Proto does a good job in this: https://capnproto.org/news/2014-06-17-capnproto-flatbuffers-sbe.html https://capnproto.org/news/2014-06-17-capnproto-flatbuffers-... "Protocol Buffers" has been the go-to for a long time but there are more options now. For uses where serialization/deserialization CPU time is a concern, it seems to really a question of Cap'n Proto versus flatbuffers ( https://google.github.io/flatbuffers/ https://google.github.io/flatbuffers/ ).
- kyrra 10y agoI'm wondering how it compares to proto3, which I understand is a fairly large change to the original protobuf.
- kentonv 10y agoProto3 is not a large change. In fact it shares most of its code with proto2, as I understand it. The underlying encoding is the same. Proto3 removes some features from proto2 which were deemed overcomplicated relative to their value (unknown field retention, non-zero default values, extensions, required fields) and adds some features that people have wanted for a long time (maps). But all of these features are things that are "on top" of the core, not really fundamental changes. I think the only change which affects my comparison post (linked by GP) is removal of unknown field retention. This is actually noted in the comparison grid. I'm honestly very surprised that they chose to remove this feature since it is critical to many parts of Google's infrastructure.
- dekhn 10y agoproto2 has maps, but like proto3 they aren't maps, they're an randomly ordered sequence of key value pairs ugh.
- kentonv 10y agoPresumably the lookup table is built at parse time? (Proto2 definitely didn't have any built-in notion of maps when I was working on it. I thought maps were added as a proto3 feature...)
- dekhn 10y agohttps://developers.google.com/protocol-buffers/docs/proto#maps https://developers.google.com/protocol-buffers/docs/proto#ma... I don't think any lookup table is provided (the wire order of entries is undefined). They are not lookup maps, they are syntactic sugar for repeated key/value pairs.
- kentonv 10y agoThe "lookup table" is constructed at parse time. Yes, the items are just key/values on the wire, but when parsed (in C++, at least) they are placed in a google::protobuf::Map, which is hashtable-based. I guess it could differ across languages -- some might not have any specific support for maps yet.
- AmrEldib 10y agoAnother (brief) overview is on Hanselminutes episode with Kenton Varda http://www.hanselminutes.com/497/your-personal-cloud-platform-with-sandstormio-and-kenton-varda http://www.hanselminutes.com/497/your-personal-cloud-platfor...
- vvanders 10y agoBig fan of flatbuffers here. Good performance, fit out needs and has MSVC pre-2015 support which is something Capt'n Proto is sorely missing.
- igrekel 10y agoWhen we had to choose a few months back, we picked Cap'n Proto over flatbuffers mainly over the fact that we liked the API better. I couldn't retrieve my notes from back then so I sadly can't point out. In any case, we never looked back.
- igrekel 10y agoWhen we had to choose a few months back, we picked Cap'n Proto over flatbuffers mainly over the fact that we liked the API better. I couldn't retrieve my notes from back then so I sadly can't point out. In any case, we never looked back.
- alfalfasprout 10y agoFor our use (serializing streaming market data from direct NASDAQ and BATS feeds) we found both Capn'Proto and Flatbuffers to perform similarly. Ultimately we went with flatbuffers because we found the API much cleaner across languages (C++/Go), but it's ultimately going to depend on your use case which one you use. Performance is pretty much identical. We populate our own custom structs from the capn'proto/flatbuffers structs anyways so we're never really doing zero-copy. That said, since both of these formats transmit numerical data as little-endian memory-representations of ints/float/longs/doubles their performance is fantastic.
- Cyph0n 10y agoFor some reason, the banner (infinitely faster?), name, and introductory FAQ-style responses made me think the whole thing is a joke - similar to Vanilla JS [1]. Anyways, it seems like a cool project, so I'll be sure to follow its development closely. [1]: http://vanilla-js.com/ http://vanilla-js.com/
- draw_down 10y agoI had the exact same reaction. I really had to read for a long time (including code) before deciding it was real. "infinitely faster" confused me too.
- kentonv 10y agoI originally released it on April 1st, 2013, with the announcement post: So, uh… I have a confession to make. I may have rewritten Protocol Buffers. Again.
- Keyframe 10y agoAdd that as a tag line. I may have rewritten Protocol Buffers, but infinitely faster.
- otoburb 10y agoAlthough this is listed on the introduction page, the author is also the same person that co-founded Sandstorm.io[1]. [1] https://sandstorm.io/about https://sandstorm.io/about
- draw_down 10y agoMmhmm.
- pjscott 10y agoThe best jokes in software usually have running code. The best of the best are practical.
- flatline 10y agoInterfaces! Inheritance! Looks promising. Protocol buffers are nice for their compact encoding and multi-language generator support but as a schema language they are really cumbersome. Composition is pretty much all you get, there are no longer required fields, you can't even use enums as a key type in a map. I'm sure their use cases are not necessarily the same as mine but sometimes I miss just using plain old XML.
- kentonv 10y agoTo be clear, Cap'n Proto's serialization layer, from a schema perspective, is almost exactly the same as Protobuf (though with a very different underlying encoding). The interfaces and inheritance relate to the RPC system. The interfaces are for remote objects. See: https://capnproto.org/rpc.html https://capnproto.org/rpc.html
- TazeTSchnitzel 10y agoBlender's file format does something similar, it essentially saves a core dump to disk.
- deleted 10y ago[deleted]
- kentonv 10y agoHi all, Cap'n Proto author here. Thanks for the post. Just wanted to note that although Cap'n Proto hasn't had a blog post or official release in a while, development is active as part of the Sandstorm project (https://sandstorm.io https://sandstorm.io). Cap'n Proto -- including the RPC system -- is used extensively in Sandstorm. Sandboxed Sandstorm apps in fact do all their communications with the outside world through a single Cap'n Proto socket (but compatibility layers on top of this allow apps to expose an HTTP server). Unfortunately I've fallen behind on doing official releases, in part because an official release means I need to test against Windows, Mac, and other "supported platforms", whereas Sandstorm only cares about Linux. Windows is especially problematic since MSVC's C++11 support is spotty (or was last I tried), so there's usually a lot of work to do to get it working. As a result Sandstorm has been building against Cap'n Proto's master branch so that we can make changes as needed for Sandstorm. I'm hoping to get some time in the next few months to go back and do a new release.
- kixpanganiban 10y agoKenton, thank you for your work on this! I currently use ZeroRPC (which uses protobuf and msgpack) and I was blown away by Cap'n Proto. Really excited to try it out soon! Some questions: - Do you guys have an RPC library written in anything other than C++? If not, could you point me to protocol specs so I can start writing my own? - Since it uses a streaming model to support random access, what encryption method do you think would work best with Cap ' n Proto that would keep it speedy and still retain all functionality? Thanks!
- kentonv 10y ago> - Do you guys have an RPC library written in anything other than C++? If not, could you point me to protocol specs so I can start writing my own? People have written implementations in Rust, Go, and Erlang, and wrappers around the C++ library in Javascript and Python: https://capnproto.org/otherlang.html https://capnproto.org/otherlang.html Scroll down that page for some info on how to start writing an implementation in another language. The RPC protocol spec is here: https://github.com/sandstorm-io/capnproto/blob/master/c++/src/capnp/rpc.capnp https://github.com/sandstorm-io/capnproto/blob/master/c++/sr... > - Since it uses a streaming model to support random access, what encryption method do you think would work best with Cap ' n Proto that would keep it speedy and still retain all functionality? Hmm, I'm not clear on what you mean by "streaming model" -- I think of "streaming" as the opposite of random access. Regarding encryption, this is a very big question and there are a lot of different needs and use cases to consider. Mostly I don't think that use of Cap'n Proto affects encryption decisions much, but if you want to make sure you don't lose random access, you should of course use a cipher that supports random access, like chacha20 or AES-CTR.
- imaginenore 10y agoIs it faster than MsgPack?
- kentonv 10y agoDepends on the use case! In fact, the answer to "is X format faster than Y format" always depends on the use case. It's always easy to construct cases where one or the other looks better. People of course want to know "on average", but in reality there's no such thing as an "average" use case. You'll ultimately have to test the case you have in mind to find out. With that said, here are some considerations: - msgpack is usually used as a binary encoding of JSON, with no schemas. That means that textual field names are included in the encoded message. Formats like Protobuf and Cap'n Proto that have schemas known in advance can avoid this bloat, making them faster and smaller. - msgpack is not a zero-copy encoding. It's necessary to parse the whole message upfront before you can use it, like with protobuf. Cap'n Proto is zero-copy, the advantages of which are described extensively on the page. For example, if you have a multi-gigabyte file containing a massive Cap'n Proto message, and you just want to read one field from one place in that message, you can do that by memory-mapping the file. No need to read it all in. That's not possible with Protobuf or Msgpack. I think it's best to focus on these kind of paradigm-shifts when trying to reason about performance. You can always micro-optimize the encoding path later on, but you can't suddenly switch to zero-copy later if your data format wasn't designed for it.
- honkhonkpants 10y agoIs X faster than Y will be tough to determine without a really detailed treatment of the use case. For example in C++ you want to account for the unavoidable construction of your class type, its inevitable destruction, how long it takes to put it on or take it off the wire, how often you might expect to miss the cache when referencing it, whether the type of movable or copyable and how expensive that is, whether or not you can reset it for reuse without destroying it, and much more. If you were to compare to protobufs I doubt that you'd find serialization and deserialization to be the predominant costs. I don't know what the cost breakdown looks like for capnproto, but maybe kentonv has standing benchmarks.
- niftich 10y agoI've always liked Cap'n Proto because it was (quite literally) the ideas behind Protobuf taken to an extreme, or, depending on your point-of-view, reduced to its most basic components: data structures already have to sit in memory looking a certain way, why can't we just squirt that on the wire instead of some fancy bespoke type-length-value struct? Of course, the hardest part is convincing everyone that it's not your bespoke type-length-value struct, but that you have good reasons for what you're doing. I think the humorous, not-so-self-serious presentation has worked in its favor (but that's just a subjective opinion and I can't back it up with data).
- makmanalp 10y agoI thought that the main reason we didn't do this was because it's was hard - platform inconsistencies like 32/64 bit, endianness, not to mention differences between how languages store things, etc. The thing that irks me about these methods is that if you're using a capnproto Int rather than a regular Int, doesn't that mean that you're basically forgoing a lot of functionality that was built around and works with the regular old data types? For example, we also do that with numpy data types in python, but there the performance benefit is super clear - numerical operations dominate. I guess it really depends on your use case. If most of your time is spend on serde, then perhaps it's worth it.
- kentonv 10y ago> platform inconsistencies like 32/64 bit, In terms of data layout, 32-bit vs. 64-bit architecture only really affects pointer size. But Cap'n Proto does not encode native pointers (that obviously wouldn't work), so this turns out not to matter. > endianness, It turns out almost everything is little-endian now. Also, big-endian architectures almost always have efficient instructions for loading little-endian data. So Cap'n Proto just has to make sure to use those instructions in the getters/setters for integer fields. > not to mention differences between how languages store things, etc. Cap'n Proto actually doesn't attempt to match how any language stores things. Instead, it defines its own layout that is appropriate for modern CPUs. It ends up being very similar to the way many languages store things (especially C), but isn't intended to exactly match. The C++ implementation of Cap'n Proto generates inline getter/setter methods that do pointer arithmetic that is equivalent to what the compiler would generate when accessing a struct. For Java, Cap'n Proto data is stored in a ByteBuffer, which effectively allows something like pointer arithmetic. Again, getters/setters are generated which use the right offsets. Most other languages end up looking like either C++ or Java.
- venning 10y agoCorrect me if I'm wrong, but some of this sounds like blitting, except optimized for the in-memory structure and not the on-disk structure. In 2008, Joel Spolsky wrote about 1990s-era Excel file formats and how they used this technique to deal with how slow computers were then [1]. Same technique, new problem set. [1] http://www.joelonsoftware.com/items/2008/02/19.html http://www.joelonsoftware.com/items/2008/02/19.html
- zaptheimpaler 10y agoCould this be used as an alternative to Apache Arrow[1]? [1] https://arrow.apache.org/ https://arrow.apache.org/
- setori88 10y agoFractalide (http://githib.com/fractalide/fractalide http://githib.com/fractalide/fractalide) is an implementation of dataflow programming (specifically flow based programming). Component build hierarchies are coordinated via the Nix package manager. Capnproto contracts are weaved into each component just before build time. These contracts are the only way compenents talk to each other. Thanks Sandstorm.io for this great software.
- edraferi 10y agoFTFY: https://github.com/fractalide/fractalide https://github.com/fractalide/fractalide
- wtbob 10y ago> The Cap’n Proto encoding is appropriate both as a data interchange format and an in-memory representation, so once your structure is built, you can simply write the bytes straight out to disk! Eh, I'd rather pay the cost of serialisation once and deserialisation once, and then access my data for as close to free as possible, rather than relying on a compiler to actually inline calls properly. > Integers use little-endian byte order because most CPUs are little-endian, and even big-endian CPUs usually have instructions for reading little-endian data. sob There are a lot of things Intel has to account for, and frankly little-endian byte order isn't the worst of them, but it's pretty rotten. Writing 'EFCDAB8967452301' for 0x0123456789ABCDEF is perverse in the extreme. Why? Why? As pragmatic design choices go, Cap'n Proto's is a good one (although it violates the standard network byte order). Intel effectively won the CPU war, and we'll never be free of the little-endian plague. It's all so depressing.
- morecoffee 10y ago> Eh, I'd rather pay the cost of serialisation once and deserialisation once, and then access my data for as close to free as possible, rather than relying on a compiler to actually inline calls properly. The thing is, as close to free as possible is surprisingly expensive. Protobuf's varint encoding is extremely branchy, and hurts performance in a datacenter environment (where bandwidth is free, and CPU is expensive). > As pragmatic design choices go, Cap'n Proto's is a good one (although it violates the standard network byte order). Intel effectively won the CPU war, and we'll never be free of the little-endian plague. Did they though? Arguably there are far more ARM CPUs (like the one in your pocket) than there are server CPUs. Since cellphones and other low power devices are almost all big endian, it seems like network byte order would have been better to use. High powered servers can pay the cost of coding them, but battery powered devices cannot afford to do so.
- niftich 10y agoARM (since v3) is bi-endian for data accesses and defaults to little. You can confirm this by searching the ARM Information Center [1] for 'support for mixed-endian data', but I can't get a working URL for that exact page. [1] http://infocenter.arm.com/help/index.jsp http://infocenter.arm.com/help/index.jsp
- morecoffee 10y ago> capability-based RPC system. This sounds like a cool idea, but so far I haven't seen any good explanation of how it works, and why it will save me from rolling my own ACL system. For bragging about it in the very first sentence, there is surprisingly little detail about how it works.
- kentonv 10y agoIt's a complicated topic -- it requires thinking about things in a different way, and tends not to make a lot of sense until at some point it "clicks" and you realize all sorts of patterns you were already using are actually special cases of capabilities. Here is some reading: https://capnproto.org/rpc.html#security https://capnproto.org/rpc.html#security https://sandstorm.io/how-it-works#capabilities https://sandstorm.io/how-it-works#capabilities http://zesty.ca/capmyths/usenix.pdf http://zesty.ca/capmyths/usenix.pdf
- sandGorgon 10y agoHow do you pronounce the name? If libreoffice is bad.. This name is absolutely impossible. Is it captain? Is it cap+n+proto? A lot of collaboration is verbal - people sit around and talk about stuff. I don't know if it is a fun take on an American word... But it is impossible to use in the rest of the world. I really wish you would call it something else... Unless it is personal for you :(
- haneefmubarak 10y agoSo yeah, in the states, we have this cereal called Cap'n Crunch that some people love. I think he was making a pun out of that. The pronunciation would thus be "cap [the sound the letter 'n' makes] crunch".
- couchand 10y agoSeems like a good opportunity to point out that the pun is that the RPC model is based on capabilities. Capabilities and Protobuf -> Cap. 'n' Proto. -> Cap'n Proto
- DonHopkins 10y agoGee, all this time I thought he was hitching his wagon to the right honorable, dapper, charismatic, articulately spoken, upstanding guru of phone phreaking culture.
- kentonv 10y agoCAP-ən PRO-to But "Captain Proto" is acceptable if you have trouble with the contraction. Or you can also think of it as "Cap-and-Proto". Which is an intentional pun ("capabilities and protocols", or something). Googling either of these will get you to the right place, so I think it gets the job done.
- sandGorgon 10y agoProtocap maybe? Btw, you rank in the 5th result for "protocap" on Google. Now that's a name all of us can pronounce!
- mixmastamyk 10y agoApache thrift doesn't seem to be mentioned, how does it compare?
- kentonv 10y agoThrift is very much equivalent to Protobuf for the purpose of everything discussed on the web site. As the site mentions, I prefer to pick on Protobuf because I am also the author of Protobuf v2. :)
- deleted 10y ago[deleted]
- Perceptes 10y agoBig fan of what Sandstorm is doing, both with Sandstorm itself and this component. I really want to use this instead of gRPC, as it seems technically superior, but language bindings and adoption across language ecosystems are likely to be a big downside given that (as Kenton mentions in a comment elsewhere here) Sandstorm isn't really interested in Cap'n Proto being widely adopted. All my new stuff is built in Rust, so the Sandstorm team's interest in and use of Rust are a good fit for me. But when it comes to interoperability with other languages, this may end up being a concern compared to gRPC. In any case, I hope to see the Rust implementation eventually replace the C++ one as the official reference implementation.
- kentonv 10y agoFWIW, language interoperability is of interest to Sandstorm since apps can be written in many languages. But we're only 7 people with a lot on our plate, so unfortunately we can't currently be the ones to go around implementing Cap'n Proto in every language. That will change as Sandstorm grows.
- Perceptes 10y agoGreat news, and thanks for the response! I don't have an immediate need for a system like this, so perhaps by the time I do there will be broader adoption.
- matmann2001 10y ago.
- megak1d 10y agoI've always liked the look of this, saw it a while back but we are still using protobuf in our .NET environment simply due to the "free" schema generation using AOP/attributes [ProtoContract]/[ProtoMember] in Marc Gravell's excellent protobuf-net (https://github.com/mgravell/protobuf-net https://github.com/mgravell/protobuf-net) project - I assume this would also be possible for cap'n proto.
- Paul_S 10y agoWould probably be a good idea to have a no-exceptions version.
- kentonv 10y agoYou can compile Cap'n Proto with -fno-exceptions, and it does a bunch of things differently to make that work. Basically, invalid-input exceptions instead replace the invalid data with a reasonable default value and set a flag on the side that you can query to see if there was any invalid input. Assertion failure exceptions (where there is no way to recover) largely turn into fatal errors.
- mrfusion 10y agoI hate to ask but can anyone explain this like I'm from 2000? I guess it's a way to send data to your front end java script but not use json and this compresses it so it's faster? How much better than using json is it?
- lmm 10y agoIt's a lot higher-performance, because you virtually don't have to parse. Remember when MySQL started supporting the memcached protocol because it'd reached the point where form simple pkey lookups it was spending more time parsing an SQL query than actually executing it? It's like doing that for your program. But for me at least, the real advantage over JSON isn't the performance but the schema compatibility. You have a spec for your data and generate code from that, which means the spec is guaranteed to be correct, and there's clear documentation about what changes to the spec are or aren't forward or backward compatibile. (You get the same thing from the original Protocol Buffers though).
- striking 10y agoIt lets you put data on the wire, in a structured format, right out of memory. Asking the question "how much faster is it" isn't even valid here, because it skips the usual serialization process. Cap'n Proto generates you some code that contains some data structures. You put data into these structures, and they will automatically be in the right shape to put directly on the wire. And then that data can be pulled right off the wire and right into memory and be fully ready to access, with no intermediate step. It's, in a sense, infinitely faster than JSON serialization or deserialization. Because it doesn't even perform any serialization. It's just data. There are some other tricks at play here, but I won't go into them. This is plenty cool.
- mrfusion 10y agoSo could django send data to jquery with this? Or what are some simple use cases?
- scott_karana 10y ago
- Twonneilb22ll 10y agoI'm happy to hear from you all in glad to see you're doing all you can for updates
- Twonneilb22ll 10y agoI'm happy to hear from you all in glad to see you're doing all you can for updates Programs codes appreciate the hard work
- Twonneilb22ll 10y agoI'm happy to hear from you all in glad to see you're doing all you can for updates Programs codes appreciate the hard work
- chaotic-good 10y agoI think that compression is a must for the serialization library. Protobuf uses almost twice less memory than Cap'n Proto. Using an external compression is not an option in some cases. E.g. consider building tcp-server that communicates with thousands of clients simultaneously. Each client connection will have its own LZ4 context that should be heap allocated. I believe it's about 16KB per connection + buffers. This results in large memory consumption and a lot of random memory access and TLB misses.
- kentonv 10y agoCap'n Proto offers "packed" encoding which applies light compression (removing zero-valued padding bytes), brings it in-line with Protobuf, and ought to be much faster than the things Protobuf does for "compression" (varint is a very slow encoding!).
- chaotic-good 10y agoI'm glad that I was wrong about it.