3 ms·
>> rusts type system is strong enough to build extremely powerful zerocopy serialization > Unclear what you mean, but other than syn and quote you don't have a
by cstrahan 2y ago
>> rusts type system is strong enough to build extremely powerful zerocopy serialization
> Unclear what you mean, but other than syn and quote you don't have a way to reflect and do code gen, outside of build script. Which also use it.
> Note there are alternative like mini serde and nano serde, but no library for type reflection.
Reflection/comptime/codegen are all orthogonal to zero copy (de)serialization.
To briefly(ish) describe zero copy by way of comparison, consider JSON: there's no way to parse a JSON object into your language's data structures without (among other things) copying any strings you come across (rather than directly borrowing a slice of the original buffer). Why not? Consider escapes: you must first unescape the string, which entails a new allocation. But this doesn't stop at strings: if you have an array, you must first parse that array (usually accumulating a copy of the elements in a Vec<Object> or similar, which in turn means that the whole array was effectively copied), rather than providing a "cursor" into the original buffer. Parsing JSON requires traversing the entire buffer and copying everything you come across.
Protocol buffers works much the same way: because structures are variable length (and in fact, scalars are too -- integers are stored as base 128 varints), if you want the Nth element of an array, you have no choice but to parse all the proceeding elements (rather than nudging a cursor's offset by N×sizeof(Elem) -- you can't do that because the size of any element is unknown until after you've parsed it). Because you want O(1) indexing after parsing, the logical implication is that whenever you parse a protobuf message, the protobuf library parses (and copies) the entire thing, and any array/repeated field ends up as a new array allocation in your language (e.g. Vec<Elem>).
Contrast with something like flatbuffers or Cap'n Proto: the code generated from your schema file gives you structures that (more or less) have two fields: a reference to a buffer, and an offset into that buffer. When you do something like person.age(), the offset of the age field (which is constant) is added to the offset of this person record, and that combined offset if used to index into the buffer (something like buffer.read_u32(offset)). Using a library like this gives you performance similar to what you'd have dereferencing array indices and struct fields in plain old data types in your language of choice. You don't pay in memory and processor time parsing everything up front, you only pay for the scalar memory reads you actually use (and a little bit of quasi-pointer chasing, not unlike what would happen with native structs).
Put another way: a zero copy (de)serialization protocol is one where the on-disk format is the same as the (readily usable) in-memory format. This rules out things like string escaping (just store the original bytes), variable sized records/integers, variable length arrays stored inline with records (otherwise that would make records themselves variable length), storing pointers in records (because those pointers will be meaningless when read from disk; instead: store offsets), etc.
You can read more about zero copy as it relates to Rust here:
https://manishearth.github.io/blog/2022/08/03/zero-copy-1-not-a-yoking-matter/ https://manishearth.github.io/blog/2022/08/03/zero-copy-1-no...
Here's the Wikipedia article:
https://en.wikipedia.org/wiki/Zero-copy https://en.wikipedia.org/wiki/Zero-copy
Examples of zero copy (de)serialization libraries:
https://github.com/rkyv/rkyv https://github.com/rkyv/rkyv
https://github.com/google/flatbuffers https://github.com/google/flatbuffers
https://github.com/capnproto/capnproto https://github.com/capnproto/capnproto
- Ygg2 2y ago> Put another way: a zero copy (de)serialization protocol is one where the on-disk format is the same as the (readily usable) in-memory format. Ok, but for it to be useful in a (Rust) program it has to be converted to a Rust type. At some point you'll have to do a conversion. Whether it's UTF-16 to String or string "false" to `bool`. The reflection, code gen, etc. is the answer how you convert the values auto-magically.
- cstrahan 2y ago> The reflection, code gen, etc. is the answer how you convert the values auto-magically. I don't think anyone (j-pb included) is saying anything to the contrary. Here's what you wrote: >> rusts type system is strong enough to build extremely powerful zerocopy serialization > Unclear what you mean, but other than syn and quote you don't have a way to reflect and do code gen, outside of build script. Which also use it. But your response doesn't logically follow from the text you quoted (so I figured you weren't familiar with zero copy). This isn't j-pb saying that Rust's type system could be used to forgo proc-macros -- j-pb isn't saying anything about proc-macros in the text you quoted. To be clear, these two points from j-pb's original comment are two separate, orthogonal issues: > - rusts type system is strong enough to build extremely powerful zerocopy serialization, but people don't explore that space because of serde > - macros and especially proc-macros in Rust are horrible, and are only feasible because of the syn crate
- Ygg2 2y ago> - rusts type system is strong enough to build extremely powerful zerocopy serialization, but people don't explore that space because of serde That's what I have a problem with. Pure Zero-copy parsers aren't explored because 99.9% of the time you have to escape and/or convert data to be useful. Let's say we create a localization library that's zero-copy. Great, now whenever we call a field, since it's zero copy and might contain an escaped value, we need to invoke the escape function on it. So every call of field actually has an overhead of a method call, know what doesn't have that overhead? Converting it once and serving it constantly. But that's not zero-copy, as per the explanation given. It's not due to serde, it's because people don't want or need a pure zero copy parser. Zero-copy parsers, as you explained, make a lot of sense if you have a packet of data, do some mapping, evaluating and send it over the wire ASAP. That's not the same use case as storing data for longer time like settings, localization, serialization, etc. > - macros and especially proc-macros in Rust are horrible, and are only feasible because of the syn crate Macros by example, don't use syn crate, they are more like regex for syntax than anything. Proc macros were always considered a kind of temporary feature that became a backbone of Rust. Pretty sure it was used a lot in some Servo components, plus it's kinda needed for many other things. Additionally, syn isn't the only way to use in proc-macro. You're 100% free to roll your own, you just have to account for all the edge cases, all minor nits in language, etc. So you can write your own buggy version of syn, or you can use syn.