10 ms·
Using Protobuf instead of JSON to communicate with a front end
- dustingetz 11y agoTransit is similar but addresses the flaws described in this article http://blog.cognitect.com/blog/2014/7/22/transit http://blog.cognitect.com/blog/2014/7/22/transit
- teh 11y agoNot sure transit is designed for the same space. E.g. there seems to be no schema, and the default JSON encoding isn't super readable either. Protobufs can be encoded as JSON and as text, so there are some ways to address the readability I guess.
- jwr 11y agoTransit is a really good solution. As for Protobuf, I tried using it in a number of places, but found it to be very inflexible (schema!) and hard to debug in case of problems.
- vruiz 11y agoI guess it only makes sense if you are already using protobuf everywhere else in your stack. Specially if you are leveraging GRPC[0] which is already profobuf over HTTP. The network tab problem could be solved by an extension, or browsers could offer the tools built-in if there were to become a trend. [0] http://www.grpc.io/ http://www.grpc.io/
- soldergenie 11y agoGrpc doesn't work over the browser. This is stated explicitly in the FAQ. See https://groups.google.com/forum/#!topic/grpc-io/5Ic8MKgltwY https://groups.google.com/forum/#!topic/grpc-io/5Ic8MKgltwY
- rektide 11y agoIt's rather painful that they don't seem to have any design docs up for their HTTP transport, leaving it to things like FAQ entries to explain these details. This was my first thought too- grpc does this.
- soldergenie 11y agoThe protocol is documented - https://github.com/grpc/grpc-common/blob/master/PROTOCOL-HTTP2.md https://github.com/grpc/grpc-common/blob/master/PROTOCOL-HTT... However, you still need the FAQ to figure out that browser transport isn't supported
- omouse 11y agoAt work we're using HTTP requests and now we added RabbitMQ in the last few months to deal with the fact that our frontend has to talk to our backend. After seeing this article it feels like we chose the wrong tool for the job; protobuf/thrift appear to be typed which would have saved us a lot of frustration as we've already run into multiple cases where the receiver or sender have messed up the type conversion or parsing.
- tokenizerrr 11y agoI don't see how protobuf is mutually exclusive with RabbitMQ. RabbitMQ is a message broker and can send around byte arrays. These byte arrays can be anything, including protobuf messages.
- kajecounterhack 11y ago(The above, but yes, send protobufs because they are typed.)
- tokenizerrr 11y agoSorry, what do you mean?
- justinsb 11y agoI like to use Protobuf in my server code, but then support JSON _or_ Protobuf as the encoding. So browsers can continue to use JSON, but the server gets strongly-typed Protobuf structures.
- PaulHoule 11y agoYeah, if you are using a statically typed language, binary formats like Protobuf are a big win, but if you are going to have the dynamic language overheard that comes with JS, there isn't much gain to be had from binary formats.
- edgarvm 11y agoWhich library do you use to support both?
- justinsb 11y agoOne I rolled myself (https://github.com/fathomdb/cloud/tree/master/fathomcloud-common/src/main/java/io/fathom/cloud/protobuf https://github.com/fathomdb/cloud/tree/master/fathomcloud-co...), but I'm switching to proto3.
- haberman 11y agoWhat you describe is exactly how proto3, the latest version of protobuf, will work! proto3 supports both binary protobuf encoding and JSON natively, so you can switch between them as desired. https://developers.google.com/protocol-buffers/docs/proto3 https://developers.google.com/protocol-buffers/docs/proto3 proto3 is currently in alpha, but we are working to bring it closer to release (I work on the protobuf team at Google).
- hesdeadjim 11y agoOh cool, didn't know there was a new version of protocol buffers. I ended up choosing Thrift for my current project due to wider language support, but I have been frustrated with some the limitations of the IDL (primarily no recursive data structures, so no generic storage of JSON-like objects).
- laurentoget 11y agoAnother way to do this is to specify the protocol in protobuf but have the server translate responses and requests to and from json. The java protobuf library does that for you out of the box. This is easier to implement. I would be curious to compare performance of both approaches in different contexts.
- deleted 11y ago[deleted]
- sbarre 11y agoThe biggest takeaway for me from this experiment was "always make sure you are gzipping your output".
- swalsh 11y agoI always wondered why google decided to build Protocol Buffers. ASN.1 seemed like it worked well, and it covered all the corners.
- VikingCoder 11y agoHere was Kenton Varda's response: https://groups.google.com/forum/#!topic/protobuf/eNAZlnPKVW4 https://groups.google.com/forum/#!topic/protobuf/eNAZlnPKVW4 My understanding of ASN.1 is that it has no affordance for forwards- and backwards-compatibility, which is critical in distributed systems where the components are constantly changing. ... OK, I looked into this again (something I do once every few years when someone points it out). ASN.1 _by default_ has no extensibility, but you can use tags, as I see you have done in your example. This should not be an option. Everything should be extensible by default, because people are very bad at predicting whether they will need to extend something later. The bigger problem with ASN.1, though, is that it is way over-complicated. It has way too many primitive types. It has options that are not needed. The encoding, even though it is binary, is much larger than protocol buffers'. The definition syntax looks nothing like modern programming languages. And worse of all, it's very hard to find good ASN.1 documentation on the web. It is also hard to draw a fair comparison without identifying a particular implementation of ASN.1 to compare against. Most implementations I've seen are rudimentary at best. They might generate some basic code, but they don't offer things like descriptors and reflection. So yeah. Basically, Protocol Buffers is a simpler, cleaner, smaller, faster, more robust, and easier-to-understand ASN.1.
- jebblue 11y ago"While I see the need for Protobuf and Thrift for services communication, I don't really see the point of using it instead of JSON for the frontend." Ah Ok whew, so the title was wrong or designed for click bait.
- ohitsdom 11y agoHas the title been updated? It currently is "Using Protobuf instead of JSON to communicate with a front end", which is not click bait at all. The author used Protobuf instead of JSON as an experiment, and concluded that there is no reason to use it.
- zubspace 11y agoOne thing, where protobuf (at least protobuf-net) really shines, is serialization of data into a binary format which is incredibly fast. In .NET, all inbuilt alternatives are slower by a large margin. https://code.google.com/p/protobuf-net/wiki/Performance https://code.google.com/p/protobuf-net/wiki/Performance
- dmsimpkins 11y agoI agree. I recently converted some large files that were previously stored using XmlSerializer to use protobuf-net, and I found an 8x increase in space efficiency, and 6-7x increase in (de)serialization efficiency. It really is a fantastic library, and if your classes are already marked up for serialization, there is very minimal work required to make the switch. For files that need not be human-readable, protobuf is definitely the way to go.
- skybrian 11y agoIt's possible to encode a protobuf as JSON and we do it all the time at Google. In browsers, native JSON parsing is very fast and the data is compressed, so going to a binary format doesn't seem worthwhile. The .proto file is used basically as an IDL from which we generate code.
- teh 11y agoCan I ask which library you are using? I found a few [1] but none seem super robust. Also, how do you deal with the bytes type? [1] https://code.google.com/p/protobuf-json/ https://code.google.com/p/protobuf-json/ https://github.com/benhodgson/protobuf-to-dict https://github.com/benhodgson/protobuf-to-dict
- philsnow 11y agoJust a guess here: I would have an agreed-upon key name suffix like "__b64_enc" or something. Serializers take a field "foo" of type bytes and serialize it as "foo__b64_enc": b64(value), and deserializers strip and base64-decode. edit: it's exactly what you would do if you wanted to pass any binary data as json over the wire, regardless of whether you're using protobufs. you'd just get it "for free" (meaning you wouldn't have to write the boilerplate, not that you don't have to en/decode).
- skybrian 11y agoThere's no standard for this and it's mostly not open source as far as I know. The overall approach is called "JSPB" but there are various flavors. A typical use case is for a web app that has its own private RPC to its own servers, so interop isn't an issue. Also, web apps generally don't need or want to work with binary data, so better not to send it. I recently became the maintainer of the Dart protobuf library which supports both JSON and binary format [1], [2]. However, the JSON format isn't necessarily compatible with other protobuf libraries you've seen. [1] https://github.com/dart-lang/dart-protobuf https://github.com/dart-lang/dart-protobuf [2] https://github.com/dart-lang/dart-protoc-plugin https://github.com/dart-lang/dart-protoc-plugin
- cletus 11y ago
- deleted 11y ago[deleted]
- nly 11y agoThrift has a JSON encoding out of the box.
- haberman 11y agoproto3 (currently in alpha) does too: https://developers.google.com/protocol-buffers/docs/proto3#json https://developers.google.com/protocol-buffers/docs/proto3#j...
- imaginenore 11y agoHave you guys tried MsgPack? If so, is it worth it? http://msgpack.org/ http://msgpack.org/
- placebo 11y agoI'm actually using it now for a project after considering various alternatives and am quite happy with the results so far.
- lux 11y agoAfter trying Protobuf on a project between C# and Javascript, we moved to MsgPack and were much happier with the ease of things.
- 7b64f0f2 11y agoUsed it to transmit data over 0MQ, worked flawlessly.
- placebo 11y agosame here :)
- w0utert 11y agoMessagePack worked great for us, fast, compact serialization, easy to use, great platform & language support. I've never used protocol buffers, mainly because I really dislike that you have to write .proto files that are then translated to code, which IMO In many situations is an unnecessary kludge. I understand it can be useful especially if you need to serialize the same things from different languages, and don't want to write the same serialization code twice (or more), but if that's not a concern for your project, I have no idea why I should prefer protocol buffers over MsgPack
- benjaminjackman 11y agoIt would probably be better to try something like Cap'n Proto or SBE if worried about performance. Otherwise I think sticking to GZIP'd json isn't going to lag that far behind. Protocol buffers biggest benefit IMHO is just their .proto file for cross language code generation. I have it on a todo list to port an SBE parser to ScalaJS. ScalaJS already backs java ByteBuffers with javascript TypedArrays. That should be really fast, the same stuff that is being worked on for making asm.js fast will also make the Cap'n Proto / SBE approach fast, so I think this has the most promise of bringing really high-performance data transfer capabilities to the browser.
- rikrassen 11y agoOne of the comments on that article was "YAY! JSON is wastefully large. I'd love to replace it." Is this true? I'm confused why JSON would be seen as a wasteful as a format. It seems to be that with any decent compression I would think it's hard to get much smaller. In this case I'm not talking about the other advantages Protobuf offers, I just want to know about size.
- maratd 11y ago> Is this true? No, it isn't true, but regardless of what format you use, there will always be someone who's not happy. Actually, I think that applies to everything in life.
- shanemhansen 11y agoThere are basically 2 areas where JSON is really wasteful. Compression can help with both of those. 1. Dictionary keys are repeated when you have an array of similar objects. 2. Non-text data. JSON can't natively represent binary data, forcing people to use things like base64 for binary and base10 for numbers.
- rikrassen 11y agoI hadn't considered binary data. Thanks.
- alkonaut 11y ago> "YAY! JSON is wastefully large. I'd love to replace it." Is this true? I'm confused why JSON would be seen as a wasteful as a format. It transmits type and field names. Depending on how complex your data is those strings could be a large part of the data. { "person": { "age": 30, "shoesize": 10 } } The above is what, 4-5 bytes of protobuf? I'm not sure what the gzipped-json data is but likely a lot more. If you were to send a list of 100 such person objects, the difference would be smaller.
- haberman 11y ago> The above is what, 4-5 bytes of protobuf? Assuming the integer fields are regular varint types (and not the "fixed" integer encoding), and assuming the tag numbers were all under 16, then this would be a six-byte protobuf.
- haberman 11y agoMaking JSON first-class is an explicit design goal of proto3, the next version of Protocol Buffers currently in alpha: https://developers.google.com/protocol-buffers/docs/proto3#json https://developers.google.com/protocol-buffers/docs/proto3#j... This will allow you to switch between JSON and protobuf binary on the wire easily, while using official protobuf client libraries. So you can choose easily whether you care more about size/speed efficiency or wire readability. Best of both worlds! I work on the protobuf team at Google and would be happy to answer any questions.
- zapov 11y agoDo you plan on improving Protobuf speed in Java? People don't expect it to be slower than JSON ;) http://hperadin.github.io/jvm-serializers-report/report.html http://hperadin.github.io/jvm-serializers-report/report.html
- haberman 11y agoI'm not super familiar with the Java implementation, but I believe it's pretty optimized and appears to do very well generally on that benchmark. One unavoidable issue is that, unlike JSON, protobuf serializers have to do two passes over the message tree, because in protobuf binary format all submessages are prefixed by their length. The first pass just calculates lengths, while the second performs the actual serialization. This could potentially slow down serialization compared to JSON, especially for message trees with lots of nodes/depth.
- abecedarius 11y agoSo encode from back to front? Then when you reach the front you know the length.
- deleted 11y ago[deleted]
- haberman 11y ago
- zapov 11y agoYou can use JSON only as codec, which can give you performance of Protobuf with much better debugability.
- oppositelock 11y agoI worked on a product inside Google which used protos (v1) as the data format to a web front end, and in practice, that system was a failure, in part to the decision to use protos. The deserialization cost of protocol buffers is too high if you're doing complex data throughput, and even though the data size is smaller, it's better to send larger gzipped JSON (which will be decompressed in native code) and deserialized into JS (also via native code). We weren't using ProtoBuf.js, but our own internal javascript implementation of a similar library, and doing all of this in JS was too expensive. Granted, we were sending around protos that had multi megabyte payloads at times. We rewrote our app eventually to send protos in JSON format to the app, while just letting our backends still pass around native protos, it worked a lot better.
- haberman 11y agoThings have changed a lot since your experience, I think. For one, a different encoding called "JSPB" has become the de facto standard for doing Protocol Buffers in JavaScript, at least inside Google. JSPB is parseable with JSON.parse(), so it avoids the speed issues you experienced. And looking forward, JavaScript parsing of protobuf binary format has gotten a lot faster, thanks in large part to newer JavaScript technologies like TypedArray. Ideally JSPB would be deprecated as a wire format in favor of fast JavaScript parsing of binary protobufs, but this would of course be contingent on the performance being acceptable. Finally, JSON is becoming a first-class citizen in proto3, so protobuf vs. JSON will no longer be an either/or, it can be a both/and. https://developers.google.com/protocol-buffers/docs/proto3#json https://developers.google.com/protocol-buffers/docs/proto3#j...
- boomzilla 11y agoWhat benefits do I get from ProtoBuf, apart from the standard binary wire format? JSON is just more popular as a serialization format. It doesn't matter what what programming language or OS I am on, there is almost always a built-in library that de/serialize JSONs at reasonable speed. To send the JSON objects around from one service to another, I can just gzip the string if it's big, or just plain UTF-8 string if it's not. ProtoBuf has to provide more values for people like me to switch. I would rather try out Apache Avro first as a replacement for what I am doing right now.
- mhahn 11y agoI'm curious if Google has a common envelope they send all service messages with. Ie. A common way of specifying pagination parameters, auth tokens etc. when sending protobuf messages between services. I've been using protobufs for my services and wrote a ServiceRequest object which has worked well. I was more just surprised about not being able to find much documentation on actual deployments as opposed to just simple tutorials.
- labianchin 11y agoI wonder how would that be like with Avro. It also has JSON encoding: https://avro.apache.org/docs/1.7.7/spec.html#json_encoding https://avro.apache.org/docs/1.7.7/spec.html#json_encoding
- krapht 11y agoHow does Protobuf compare with Corba? I'd be interested in anybody's experience if they have used both.
- nostrademons 11y agoCORBA was ridiculously complex, because they tried to make remote objects look like local ones, with messages, reference counting, naming, discovery, etc. Protobuf is just a serialization mechanism. You're thinking at a lower level of abstraction - it's all just PODs that go over the wire, you build your own RPC framework on top of that (or use gRPC, which is Google's protobuf-over-HTTP2 RPC library) and think in terms of requests & responses. IMHO trying to make everything look like an object was a mistake, and newer RPC frameworks like gRPC, Thrift, and JSON-over-HTTP are much easier to use than the late-90s frameworks like RMI, CORBA, and DCOM. Sometimes you don't want abstraction, because it abstracts away details you absolutely need to think about.
- flavor8 11y ago> Reading time: ~15 minutes. 842 words including code. Average adult reading speed: 300 words/minute. Does not compute.
- Keats 11y agoI know, I included some time for people wanting to to open some links, the github project etc. Only reading the text itself takes indeed less than 5 minutes, not sure which approach people prefer.
- rqebmm 11y agoHaving used both on a few projects, including a JS frontend, my advice is: "Don't use protobufs if you don't have to". Protobufs can be much faster, and provide a strict schema, but it comes at the price of higher maintenance costs. JSON is much simpler, easier to implement, and MUCH easier to debug. If your GPB looks like it's building properly, but fails to parse, it's a huge pain to try and decode/debug the binary. You'll wish you could just print the JSON string. If you need the speed and schema, then GPBs are great. In our case, we got a huge speed boost just by avoiding string building/parsing inherent in JSON.
- nfmangano 11y agoCould you elaborate on the maintenance costs? We use ProtoBufjs for our own real-time whiteboarding webapp over web sockets, and in the long run having strict schemas has saved us a lot of time. We're a distributed team with different members working on the front and backends, and we frequently refer to our proto files to remember how data is transferred and how it should be interpreted (explained in our proto commented code). Are the maintenance costs related to debugging unparsable messages? We've almost never had an issue there, so maybe we've just been lucky?
- sdenton4 11y agoIn my experience, it's not that tough to write a 'proto-to-dict' function in python, which lets you crack open the proto and look at its juicy innards...
- wora 11y agoCo-author of proto3 here. Proto3 was specifically designed to make proto more friendly in variety of environments, which includes native JSON support. New Google REST APIs are defined in proto3, which are open sourced[1]. [1] https://github.com/google/googleapis https://github.com/google/googleapis
- gobengo 11y ago+1 to "Did this in a real product and fully regret it"
- drawkbox 11y agoThere is definitely a place for binary serialization/de-serialization and transmission. Inter-system communication is probably the best place for binary or any place that needs high speed real-time communication with the smallest size to fit in MTU limits (game protocols over UDP for instance). Any place that you control the client and server is ok to use binary. However, I do feel there is a strange swaying back to binary (Protobuf/HTTP/2/etc). Developers are trying to wedge it in now in places it may cause more problems because it is more efficient in performance but not in use or implementation. Plus, like mentioned in this thread, you can compress JSON to be very small to send over the wire which makes the compactness of it a non-issue in non real-time cases. Going binary just to go binary is more trouble than it is worth in most cases. - Binary over keyed plain text (JSON) is harder to generically parse objects i.e. dictionaries/lists for just a few fields/keys. - Binary over JSON also seems to lock down messaging more, people have more work to change binary explicit messages because of offset issues and client/server tools must be in sync rather than just adding a new key that can be pulled as needed. - Third party implementation and parsing of JSON/XML is more forgiving making version upgrades and changes easier to do. This is especially apparent on projects that are taken over by other developers. - The language/platform on the backend leaks into the messaging. For instance Protobuf only runs on js/python currently and has various versions. The best messaging is independent of the platform and versioning is easier. I would bet binary formats end up causing more bugs over keyed/plaintext (JSON/XML and possibly compressed) though I have nothing to back that up by except my own experience largely in game development where networking state is almost always binary, for server/data I wouldn't use it unless it needs to be real-time. That being said Protobuf is awesome and I hope developers are using it where it is best suited and that developers don't start obfuscating messaging for performance where it doesn't really need to be, better to be simple unless you need to make it more complex at every level.
- Animats 11y agoWith one end in Python 2 and the other end in Javascript, using binary protobufs seems misplaced optimization. It's nice to know the support is there (well, not in Python 3, apparently), in case you need to talk to something that speaks protobufs. I'm looking forward to seeing protobufs in Rust as a macro. It should be possible; there's an entire regular expression compiler for Rust as a compile-time macro, which is a useful optimization.
- kibwen 11y agoIt's not protobuf, but Rust has quite good Cap'n Proto support: https://crates.io/crates/capnp https://crates.io/crates/capnp