7 ms·
This is going to sound sarcastic but it's not: Can we get back to just putting the members of C structures into network byte order and sending that over the wi
by doublement 7y ago
This is going to sound sarcastic but it's not: Can we get back to just putting the members of C structures into network byte order and sending that over the wire in binary, à la 1995?
- mikece 7y agoIs that less of a configuration mess than WCF was? JSON isn't "The Magical Elixir" of data exchange and I'm more than open to something better but at least we (in the .NET community) have moved past the WCF configuration nightmares.
- bob1029 7y agoWCF is an unmitigated dumpster fire. We have actually written a non-WCF client that uses a raw HttpClient implementation with StringBuilder to compose SOAP envelopes around cached XMLSerializers in order to talk to other WCF services. First request delay went from 1-2 seconds down to a few milliseconds. Memory overhead is negligible now. Prior, you could watch task manager and immediately recognize when WCF is "warming up". Additionally, the XML serializer in .NET seems almost pathologically determined to ruin everything you seek to accomplish. By comparison, JSON contracts are an absolute joy to work with. We still practice strong-typing on both sides of the wire (we control both ends), and have pretty much nothing to complain about. If you are concerned with space overhead w/ JSON, simply passing it through gzip can get you down to a very reasonable place for 99% of use cases. I understand that there are arguments to be made against JSON for extremely performance sensitive applications, but I would counter-argue that these are extremely rare in practice.
- heavenlyblue 7y agoIsn’t the problem with WCF rather than XML?
- btown 7y agoThis is more or less what https://capnproto.org/ https://capnproto.org/ does.
- omginternets 7y agoFrom the capnproto docs: >Isn’t this all horribly insecure? >No no no! To be clear, we’re NOT just casting a buffer pointer to a struct pointer and calling it a day. Isn't this a direct contradiction to your claim? Or have I misunderstood them?
- deleted 7y ago[deleted]
- lidHanteyk 7y agoCapn is better than C at struct layout. We are not, under any circumstances, going back to the 90s. We are moving forward and learning from mistakes.
- jtolmar 7y agoIIRC: capnproto generates messages that you could deserialize by casting them to the right struct, but refrains from actually doing it that way. Instead it generates a bunch of accessor methods that parse the data, as if you were reading something that's not basically a c-struct, like a protobuff.
- kentonv 7y agoThat's basically correct. Cap'n Proto generates classes with inline accessor methods that do roughly the same pointer arithmetic that the compiler would generate for struct access. There's a couple subtle differences: * The struct is allowed to be shorter than expected, in which case fields past the end are assumed to have their schema-defined default values. This is what allows you to add new fields over time while remaining forwards- and backwards-compatible. * Pointers are in a non-native format. They are offset-based (rather than absolute) and contain some extra type information (such as the size of the target, needed for the previous point). Following a pointer requires validating it for security. (Disclosure: I'm the author of Cap'n Proto.)
- asveikau 7y agoRe-read the comment I think. It doesn't say casting a struct pointer. It says putting the members of the struct into network byte order over the wire. I read that as individually serializing each member in a portable, safe way. Anyway even if you do choose the struct pointer hack (which I do not see advocated here) it can be done relatively well albeit requiring language extensions and a bit of care. Pragmas and attributes to ensure zero padding and alignment between members. No pointer members. Checking sizes and offsets after a read (the hardest part).
- izacus 7y agoThis is practically exactly what Protobuffers are. Except that they actually are defined clearly enough for multiple services written in multiple languages can work with them.
- cma 7y agoI thought that's what capt'n proto was, not protobuffers.
- CoolGuySteve 7y agoDefinitely not, protobuf's strange wire format becomes apparent if you ever look at the hexdump of one or the profiler output of your favourite protobuffer-decoding C/C++ application. They're actually kind of performance heavy for no benefit.
- imtringued 7y agoI have once looked at a benchmark that compared protobuffer, message pack, json and a variety of other serialization formats. In terms of reducing bytes per message gzipped json was ahead of all of them at the cost of increased CPU time for gzip. Protobuffer did pretty poorly, the only benefit was decreased CPU usage. I'm sure you could use some other compression algorithm like LZMA to get both good compression and good performance for JSON messages.
- CoolGuySteve 7y agoI use LZ4 (with "best" compression) for packet captures and replay with great results. I get about a 37% compression ratio with extremely fast decoding, like 10 million packets per second off an SSD. It was better than snappy, gzip, and bz2 for the trade-off of compression time, decompression time and file size. As for protobuf: flatbuffers, capn proto, HDF5, and plain C structs all deliver much, much faster decoding time. It's really not the best answer for any serialization at this point but it's still inexplicably popular.
- touisteur 7y ago
- amacbride 7y agoXDR FTW!
- pantalaimon 7y agoThis gets hairy the moment you want to add new fields.
- CoolGuySteve 7y agoBoth protobuf and plain C structs are append only formats if you put the message-type and size at the start of the C-struct.
- marcan_42 7y agoC structs do not compose extensively. Protobufs do. You can't put variable length data into a struct, and hence you can't put extensible structs into it either.
- doublement 7y agoYou can put variable length data into a struct: https://en.wikipedia.org/wiki/Flexible_array_member https://en.wikipedia.org/wiki/Flexible_array_member
- xyzzyz 7y agoOnly one field can be variable length, and it must be last. I'll pass.
- CoolGuySteve 7y agoYou definitely can but it's not as obvious, make a separate message type for list elements and append them on the wire. If you only have one list at the tail, you can use a flexible array[] at the end but it's finicky to deal with if you need more than one. You can build large hierarchical structures of messages with lists contained therein. It's pretty much how .mov/.mp4/many, many media container formats work. The technique dates back to the Amiga days.
- doublement 7y agoYeah I know that pain, there needs to be a consistent header with a version in all the messages.
- BubRoss 7y agoWhy even put them in network byte order? Every modern system is little endian, if you standardize on that, only exotic systems would have to deserialize anything.
- doublement 7y agoBecause when someone builds a hugely popular exotic system in the future, because it is one (1) cent cheaper, you'd end up with code that has to check to see if it's running on such a system.
- BubRoss 7y agoThis doesn't make any sense for multiple reasons, but especially because you wouldn't be checking anything in the first place. A big endian system would would reorder bytes and a little endian system would just use it directly from memory without another copy or reordering anything.
- bserfaty 7y agoHa - I just had this exact argument yesterday. Why indeed?
- toast0 7y agoThere's not a library pattern for host to little endian, or little endian to host, like we have with hton and ntoh. Which makes it more likely to be messed up.
- syncsynchalt 7y agoIf you force the most common system to translate byte order, then you'll have some confidence that your code is performing the translation correctly. If instead you rely on hoping that everyone added the correct no-op translation calls everywhere, you'll find your code doesn't work as soon as you port it to another CPU. This is a nice side effect of network byte order being the opposite of the dominant cpu order, though obviously it was never intended.
- nabla9 7y agoThe memory layout of a C struct is ABI and compiler dependent. Some compilers conform to same ABI in same system or similar system and work almost exactly the same, so you may grow old thinking that's how it is until it's too late. I think gcc, clang and Intel work almost the same in Linux and OSX.
- doublement 7y agoIndeed, that's why I specified putting the members of the C structure on the wire, not the structure as a whole, so it's just basic types in network byte order (i.e. consistent endian-ness) being sent.
- itronitron 7y agoI've worked on an application where that was the standard data transfer scheme, and then while working with protobuf on another project felt that after looking under protobuf's covers it was doing something very similar but wrapping an entire API around it.
- CoolGuySteve 7y agoNo, not really. #pragma pack and/or __attribute__((packed)) have been supported for eons now and guarantee the alignment of struct members between compilers. In newer C++ specs, you can also static assert that the struct is a POD type to statically ensure that there's no accidental vtable pointer. This argument pops up every time someone mentions this and every time it's completely uninformed.
- cyphar 7y agoThough it should be noted that packed structures cause compilers to produce absolutely garbage code when accessing them (because most of the accesses become unaligned) and it becomes incredibly memory-unsafe (as in "your program may crash or corrupt memory") to take pointers of fields inside the struct because they are (usually) presumed to be aligned by the compiler. Explicit alignment doesn't suffer from this problem nearly as badly (yeah, you might have to add some padding but that's hardly the end of the world -- and if you have explicit padding fields you can reuse them in the future).
- innagadadavida 7y agoProtobufs have a pretty nice variable encoding integer wire format. This gives you the flexibility of saving space without doing compression. While zero copy is nice, you cannot make it work when using compression.
- caffeine 7y agohttps://github.com/real-logic/simple-binary-encoding https://github.com/real-logic/simple-binary-encoding is a good way to do this
- ping_pong 7y agoXDR and ONC/RPC for the win!
- salgernon 7y agoI would like to submit apples archaic “Rez”[1] as a great language for declaring binary formats. It was designed to be able to describe c and pascal structures. [1] http://preserve.mactech.com/articles/mactech/Vol.14/14.09/RezIsYourFriend/index.html http://preserve.mactech.com/articles/mactech/Vol.14/14.09/Re...
- rbanffy 7y ago> Can we get back to just putting the members of C structures into network byte order and sending that over the wire in binary, à la 1995? I hope not (I know you are being sarcastic). We should use something that's trivial to implement correctly, as well as easy to read and to debug.
- sagarm 7y agoThe wire encoding for protos is much more compact than the in-memory representation, especially for sparsely populated messages (very common especially in mature systems). You'd still have to figure out some way to serialize nested messages. Note that you can have recursive message definitions.