5 ms·
Signing data structures the wrong way
- Retr0id 6mo agoPutting domain separators in the IDL is interesting but you can also avoid the problem by putting the domain separators in-band (e.g. in some kind of "type" field that is always present). Tangentially, depending on what your input and data model look like, canonicalisation takes O(nlogn) time (i.e. the cost of sorting your fields). Here I describe an alternative approach that produces deterministic hashes without a distinct canonicalization step, using multiset hashing: https://www.da.vidbuchanan.co.uk/blog/signing-json.html https://www.da.vidbuchanan.co.uk/blog/signing-json.html
- majormajor 6mo agoI think a lot of people assume that the "name" of the type, for protos, will be preserved somewhere in the output such that a TreeRoot couldn't be re-used as a KeyRevoke. It makes sense that it isn't - you generally don't want to send that name every time - but it's non-obvious to people with a object-oriented-language background who just think "ah, different types are obviously different types." The serialization cost objection is generally what I've often seen against in-bound type fields and such, as well, so having a unique identifier that gets used just for signature computation is clever. What's over my head possibly, from skimming it, about your multiset hashing is how it avoids the "these payloads have the same shape, so one could be re-sent as the other" issue? It seems like a solution to a different problem?
- Retr0id 6mo agoMultiset hashing is not related to the domain separation problem, but it is related to the broader "signing data structures" problem. (I realise my comment reads a bit unclearly, it's basically two separate comments, split after the first paragraph)
- kccqzy 6mo agoThis is just a mismatch between nominal typing and structural typing. Protobuf is basically structural typing. You can serialize a message defined with one schema and deserialize the result to a message with a different schema if the two schemata are compatible enough. Almost all normal programming languages use nominal typing. If you have `struct A {int a; int b};` it is distinct from `struct B {int a; int b};`.
- actionfromafar 6mo agoC does too as a language, but it’s fairly easy to slip up at link time or runtime. At some point the types melt away and you sit there with pointers and offsets. Again, it’s not strictly the language’s fault (I think, I’m far from a standards lawyer).
- cousin_it 6mo agoI think it's nice to be able to do things like rename nested structs and keep wire compatibility when upgrading two parts of the system at different schedules. Protos are neat. Think like a proto. (Not saying the signing problem in OP is invalid of course. Just a different problem.)
- Already__Taken 6mo agoThats a fun read and a topic I've specifically looked into before and its one of those black holes of terms than cannot be searched. You get so much noise.
- tantalor 6mo agoSince the example was given in proto, I'll suggest a solution in proto: add a message option. extend google.protobuf.MessageOptions { optional uint64 domain_separator = 1234; } message TreeRoot { option (domain_separator) = 4567; ... }
- hrmtst93837 6mo ago[flagged]
- formerly_proven 6mo agoThis article claims that these are somewhat open questions, but they're not and have not been for a long time. #1 You sign a blob and you don't touch it before verifying the signature (aka "The Cryptographic Doom Principle") #2 Signatures are bound to a context which is _not_ transmitted but used for deriving the key or mixed into the MAC or what have you. This is called the Horton principle. It ensures that signer/verifier must cryptographically agree on which context the message is intended for. You essentially cannot implement this incorrectly because if you do, all signatures will fail to verify. The article actually proposes to violate principle #2 (by embedding some magic numbers into the protocol headers and presuming that someone will check them), which is an incorrect design and will result in bad things if history is any indication. Principles #1 and #2 are well-established cryptographic design principles for just a handful of decades each.
- ahtihn 6mo agoMaybe I'm misunderstanding the article but I'm fairly sure the magic number is not transmitted. It's used exactly as you say: a shared context used as input for the signature that is not transmitted.
- lokar 6mo agoNo, I'm pretty sure they are saying you need to transmit it
- nightpool 6mo agoNo, they propose just concatenating it with the data received from the network > it makes a concatenation of the domain separator (@0x92880d38b74de9fb) and the serialization of the object, and then feeds the byte stream into the signing primitive. Similarly, verification of an object verifies this same reconstructed concatenation against the supplied signature. > Note that the domain separator does not appear in the eventual serialization (which would waste bytes), since both signer and receiver agree on it via this shared protocol specification. Encrypt, HMAC, and hash work the same way
- Muromec 6mo agoSo another lesson had been relearned from asn.1. I'm proud of working in this industry again! Next we will figure out to always put versions into the data too
- jbmsf 6mo agoThat was my first thought as well.
- maxtaco 6mo agoI would say two problems with the asn.1 approach are: (1) it seems like too much cognitive overload for the OIDs to have semantic meaning, and it invites accidental reuse; I think it matters way more that the OIDs are unique, which randomness gets you without much effort; and (2) the OIDs aren't always serialized first, they are allowed to be inside the message, and there are failures that have resulted (https://nvd.nist.gov/vuln/detail/cve-2022-24771 https://nvd.nist.gov/vuln/detail/cve-2022-24771, https://nvd.nist.gov/vuln/detail/CVE-2025-12816 https://nvd.nist.gov/vuln/detail/CVE-2025-12816) (edit on where the OIDs can be, and added another CVE)
- themafia 6mo agoThose CVEs seem a little more subtle than OID serialization issues. In the first example there are actually two distinct problems in concert that lead to the vulnerability, one of which is when a "low public exponent" is used. https://github.com/digitalbazaar/forge/commit/3f0b49a0573ef1bb7af7f5673c0cfebf00424df1 https://github.com/digitalbazaar/forge/commit/3f0b49a0573ef1...
- logicallee 6mo agoalong the same lines, did you know that you can get an authenticated email that the listed sender never sent to you? If the third party can get a server to send it to themselves (for example Google forms will send them an email with the contents that they want) they can then forward it to you while spoofing the from: field as Google.com in this example, and it will appear in your inbox from the "sender" (Google.com) and appear as fully authenticated - even though Google never actually sent you that. This is another example where you would think that "who it's for" is something the sender would sign but nope!
- tennysont 6mo agoI asked about this on the PGP mailing list at one point, and I think I was told that the best solution is to start emails with "Hi <recipient>," which seems like a funny low-tech solution to a (sad) problem.
- HanyouHottie 6mo agoThe solution to this problem without needing to modify your message is to use a protocol that will sign, then encrypt, then sign again. See section 5 here [1] or section 15 here [2]. [1] https://theworld.com/~dtd/sign_encrypt/sign_encrypt7.html https://theworld.com/~dtd/sign_encrypt/sign_encrypt7.html [2] https://computerresearch.org/index.php/computer/article/view/100954/2-High-Security-by-using-Triple_html https://computerresearch.org/index.php/computer/article/view...
- tennysont 6mo ago> without needing to modify your message Careful. I argue this is even worse. In this convention, you need to change the behavior of others. If I send a message to Alice with contents "Hey, I can't meet today" using your sign-encrypt-sign scheme, then Alice can take the inner most layer and use it to impersonate me. Alice can send "Hey, I can't meet today" to Bob at any time. I must rely on Bob demanding proof that he was, in fact, the intended recipient. From the first link: > Note though that an effective security standard should require not only that the author must provide one of these five proofs, but also that the recipient must demand some such proof as well. If your convention was upgraded into a protocol with automatic verification, then that would be different.
- deleted 6mo ago[deleted]
- deleted 6mo ago[deleted]
- jeffrallen 6mo agoThis is a nice explanation of an obvious idea. Both domain separation, and putting the domain signifier into the IDL are fine, but not novel. Crypto is hard. Do it right. Get help from your tools. 'Nuff said. Jeeze, I'm getting too old for this crap.
- lukev 6mo agoSo, isn't this a rather longwinded way to say that a signature only extends to the scope of the message it contains? It doesn't matter if I sign the word "yes", if you don't know what question is being asked. The signature needs to included the necessary context for the signature to be meaningful. Lots of ways of doing that, and you definitely need to be thoughtful about redundant data and storage overhead, but the concept isn't tricky.
- maxtaco 6mo agoHi, post author here. Agree that the idea isn't tricky, but it seems like many systems still get it wrong, and there wasn't an available system that had all the necessary features. I've tried many of them over the years -- XDR, JSON, Msgpack, Protobufs. When I sat down to write FOKS using protobufs, I found myself writing down "Context Strings" in a separate text file. There was no place for them to go in the IDL. I had worked on other systems where the same strategy was employed. I got to thinking, whenever you need to write down important program details in something that isn't compiled into the program (in this case, the list of "context strings"), you are inviting potentially serious bugs due to the code and documentation drifting apart, and it means the libraries or tools are inadequate. I think this system is nice because it gives you compile-time guarantees that you can't sign without a domain separator, and you can't reuse a domain separator by accident. Also, I like the idea of generating these things randomly, since it's faster and scales better than any other alternative I could think of. And it even scales into some world where lots of different projects are using this system and sharing the same private keys (not a very likely world, I grant you).
- cogman10 6mo agoWhy not digest the type as part of the hash? This avoids the problem in the article and keeps the transmission size small.
- maxtaco 6mo agoIt should be possible to change the name of the type, and this happens often in practice. But type renames shouldn't break preexisting signatures. In this scheme you are free change the type name, and preexisting signatures still verify with new code -- of course as long as you never change the domain separator, which you never should do. Also you'd need to worry about two different projects reusing the same type name. Lastly, the transmission size in this scheme remains unaffected since the domain separators do not appear in the serialized data. Rather, both sides agree on it via the protocol specification.
- actionfromafar 6mo agoThat’s easily addressed. We just need a global immutable registry of types, their names, an alias list and revocation list. ;-) We can let one be managed by ICANN and the others various competing offerings on ETH.
- tennysont 6mo agoThey use a magic number, rather than a digest derived from the schema[1], but otherwise they do as you suggest. The magic number is given to the signing function (sender side) and the validation function (receiver side) but does not increase the size of the transmitted message. [1] I think that's what you mean by digest, but maybe you just mean `type` = `magic number`
- patrakov 6mo agoThe problem here is that a digest derived from the schema would just reintroduce the possibility of confusion of identically encoded but semantically distinct types.
- efitz 6mo agoWhen my data structures are messages to be sent over a network, I always start with msgId and msgLen, both fixed width fields. This solves the message differentiation problem explicitly, makes security and memory management easier, and reduces routing to: switch(msg.msgId): …
- i386 6mo agoswitch on version, then messageId…
- volume_tech 6mo ago[flagged]
- colek42 6mo agoDSSE is great for this, if you need more schema use in-toto
- bengt_trustpay 6mo ago[flagged]
- socketcluster 6mo agoThe crypto dev community has a strange idea that working with binary is superior. For many algorithms, it's not. It just obfuscates what's happening and the performance advantage is negligible... Especially in the context of all the other logic in the system which uses far more resources. I didn't know that Protobuf wasn't canonical but even without this knowledge, there are many other factors which make it an inferior format to JSON. Also, on a related topic; it seems unwise that essentially all the cryptographic primitives that everyone is using are often distributed as compiled binaries. I cannot think of anything more antithetical to security than that. I implemented my own stateful signature algorithm for my blockchain project from scratch using utf8 as the base format and HMAC-SHA256 for key derivation. It makes it so much easier to understand and implement correctly. It uses Lamport OTS with Merkel MSS. The whole thing including all dependencies is like 4000 lines of easy-to-read JavaScript code. About 300 lines of code for MSS and 300 lines for Lamport OTS... The rest are just generic utility functions. You don't need to trust anyone else to "do it right" when the logic is simple and you can read it and verify it yourself! Simplicity of implementation and verification of the code is a critical feature IMO. If your perfect crypto library is so complex that only 10 people in the world can understand it, that's not very secure! There is massive centralization and supply chain risk. You're hoping that some of these 10 people will regularly review the code and dependencies... Will they? Can you even trust them? Choosing to use a popular cryptographic library which distributes binaries is basically trading off the risk of implementation mistake for the risk of supply chain attack... Which seems like a greater risk. Anyway it's kind of wild to now be reading this and seeing people finally coming round to this approach. I've been saying this for years. You can check out https://www.npmjs.com/package/lite-merkle https://www.npmjs.com/package/lite-merkle feedback welcome.
- i386 6mo ago> I didn't know that Protobuf wasn't canonical but even without this knowledge, there are many other factors which make it an inferior format to JSON. If strict typing and versioning are useless to you, sure.
- thunderfork 6mo ago
- erpellan 6mo agoAm I missing something or would this be solved by adding a 1 byte `msg` field to the payload?
- Someone 6mo agoTLDR: the idea is - to have a convention to, instead of signing “payloads, to always sign “type identifier + payload”, to prevent adversaries from reusing your signature to sign the same payload, interpreted as a different type. - use 64-bit type identifiers - put the identifiers in the IDL (may need augmenting IDL to allow that) #1 makes sense to me; #3 also makes sense, as that’s the place where people will have to look to learn about your types. #2, I think, is up for discussion. These could be longer, Java-like strings “com.example.Foo”, or whatever. I think some people also may disagree with the argument that putting type identifiers inside the payload makes messages too large, but I don’t have enough experience on that to make a judgment.