25 ms·
The specs behind the specs – a deep-dive on ASN.1
- dingosity 5y agoWhy are we even talking about ASN.1/DER/BER? We should, like the ancient Egyptian priests who opposed Akhenaten, chisel it's name from every public edifice. Referring to it not as "ASN.1, the platform-independent abstract type system," but "the great heresy, which shall not be named."
- ggm 5y agoHand in your X.509 certificate at the door. (There was a proposal to do x509 as s-expressions. The road not taken...) ANY DEFINED BY ANY -- all following comments
- cryptonector 5y agoASN.1 is glorified S-expressions. BER/DER/CER is binary S-expressions.
- ggm 5y agoAnything you can represent in a parse tree is an s-expression the point was during the discussions of SPKI we talked about using canonical s-expression notation rather than ASN.1 to represent the TBS and other structural forms. SPKI used them. It stood in contradistinction to X.509 in ASN.1 If my memory serves me right Kent and others argued for continuance of ASN.1 to keep alignment with the CCITT/ITU on standards docs. This was a long time ago. Around rfc3270 days so early 2000s. My memory is hazy and my email archive off-line. All objects in RPKI are in ASN.1 as are SNMP. I only have to deal with the former these days.
- cryptonector 5y agoSo, my take is that depending on canonical encodings in security protocols is a mistake. What one ends up doing is something like: - "hmm, I've a decoded struct here, and a signature of its original encoding, and I have to validate that signature somehow... what do I do??" - and then "ah, I know! I'll re-encode that struct and then I can validate the signature!!1!", - but now you need a canonical encoding ruleset, otherwise if the signer had any liberties at all in the encoding, you will have interoperability problems! And it turns out that specifying and -worse- implementing canonical encodings can be hard. Think of a canonical JSON... Let's say you have a JSON encoder lying around, and now you need to make it emit canonical JSON. You start by eliminating interstitial whitespace and you are ready to declare victory when you notice that you still need a canonical encoding of numbers, and also strings! Ok, now you have less-obvious design choices to make. Worse, adjusting your floating point number printer to emit canonical numbers turns out to be really hard, and there are a lot of traps in doing that. So maybe you decide you're going to limit yourself to integers. And it's all like this. There is a better answer. The Heimdal ASN.1 compiler has a --preserve-binary=TYPE option where you can say that you want the decoder to preserve the original encoding of the give TYPE(s) so that you can validate signatures later. The way this works is that for each such TYPE, the compiler adds a `_save` field that has a copy of the encoding of that type as it was seen by the decoder. I'm with Stephen Kent on this. I don't like the OpenSSH certificate format, for example -- it's missing important things and it's not that much simpler than the PKIX certificate format. The OpenSSH certificate format is much less bloaty than the PKIX one because PKIX uses DER and OpenSSH doesn't -- but so what, one could simply use an OER encoding of PKIX certificates and get the same de-bloating benefit with much less churn to existing codebases.
- pphysch 5y ago> You might have heard of similar such abstract syntax notations used for interface definitions such as Google Protocol Buffers, or Facebook’s Apache Thrift, but those languages have not been managed by a standardization organization, so the owning corporations could (in theory) make breaking changes or change the license or even remove the language definitions overnight. Is this really the main difference between ASN.1 and Google protobufs, that one is managed by a private corporation and the other by a standardization organization? Can they otherwise be used "interchangably" in designing interfaces, a la two different programming languages (with different syntax of course)?
- jwalton 5y agoIn terms of tooling, there’s excellent tooling for ASN.1 for C and C++ and maybe some other languages. There’s excellent tooling for protobufs for a handful of languages too, but they’re different sets, so in practice what languages you want to use would likely come into play.
- breser 5y agoHow excellent the ASN.1 tooling is depends on which subset of ASN.1 you're using. Some of the tooling supports one iteration of ASN.1 or the other. To the degree that the IETF had to write a document on how to deal with this since some of the standards use the older ASN.1 and some use the newer ASN.1: https://tools.ietf.org/id/draft-ietf-pkix-asn1-translation-03.html https://tools.ietf.org/id/draft-ietf-pkix-asn1-translation-0... Interoperability with ASN.1 is very fragile at best.
- cryptonector 5y agoWe have tons of interoperable PKIX implementations (OpenSSL and derivatives, NSS, OpenJDK's, GnuTLS, wolfSSL, Heimdal, and many many more), and a bunch of interoperable Kerberos implementations (MIT Kerberos, Heimdal, Windows / AD, OpenJDK's, the IBM Java's, GNU Shishi, there's a python implementation).
- cryptonector 5y ago
- AceJohnny2 5y agoCan the veterans of the 90s SSL Wars explain the issues with ASN1/DER/BER? Looking it up today, it seems like a pretty smart and extensive serialization system, and I have to wonder why new systems like Google Protobufs chose to reinvent the wheel. Conversely, how have modern systems avoided the pitfalls (if any) of ASN1/DER/BER?
- jwalton 5y agoWhen I first saw protobufs, I wondered exactly the same thing. There’s an “XER” if you want a human-readable XML encoding, too.
- carapace 5y agoPersonally, I think that people just like to reinvent things. I don't want to sound shitty (or have kentonv show up again to scold me for it) but I get the feeling that, a lot of the time, it's just that simple. https://news.ycombinator.com/item?id=20725550 https://news.ycombinator.com/item?id=20725550
- juanbyrge 5y agoTo me that is a specious argument. It's like asking why Python was invented when Cobol could suffice. The dozens of ASN.1 specs are absolutely hideous and entrenched in obsolete telecom jargon. If the sole goal Protobuf was to avoid having Google engineers be required to refer to the dozens of ASN.1 specs when disagreements or confusions arose, then it would have been 100% worth it for just that reason.
- carapace 5y agoFirst, let me confess that I don't have enough experience with ASN.1 or Protobufs to have an informed opinion. The supporting argument for the "because it's there" hypothesis for why people reinvent things (in IT) is that they do it so often. Even if all the newer message/serialization systems are better than ASN.1, they're not all better than each other, eh? Why so many? Same goes for chat systems, programming languages, etc.
- deleted 5y ago[deleted]
- andrewmcwatters 5y agoWhat’s so great about ASN.1 and it’s encoding rules is that anyone writing type-length-value serialization for networking purposes, for example[1], is basically independently reinventing ASN.1 because it’s so fundamentally optimal. It truly will make you wonder why Protobufs and others exist. [1]: https://github.com/Planimeter/grid-sdk/blob/master/engine/shared/typelenvalues.lua#L199 https://github.com/Planimeter/grid-sdk/blob/master/engine/sh...
- faebi 5y ago[2]: https://github.com/openssl/openssl/issues/4320 https://github.com/openssl/openssl/issues/4320 I recently had to deal with that situation too.
- memling 5y ago> What’s so great about ASN.1 and it’s encoding rules is that anyone writing type-length-value serialization for networking purposes, for example[1], is basically independently reinventing ASN.1 because it’s so fundamentally optimal. The challenge arises if you have very large values: by nature, TLVs require that the V be encoded before you can plug in the L. If you use definite-length encodings (as required by DER), you may end up having to hold and encode a pretty large piece of data in memory. You can work around this, of course, but it can be a challenge. Tags in ASN.1 as noted in another comment can also be pretty complicated: there are four tagging classes, and tags can be applied implicitly, explicitly, or automatically depending on the specification. This can make life a bit uncomfortable at times. On the balance, I can understand why people find ASN.1 such a pain, especially if you're not inclined to fork over money to have someone else deal with the encodings. For medium- to large-sized companies, though, it's probably not a bad deal: get a support contract from one of the commercial vendors, get training, and save yourself six man-months on writing pretty bullet-proof serialization code without the headache of worrying about standards incompatibilities. If you happen to work in telecommunications or security, you're going to deal with ASN.1 at some point anyway, so having something that can talk to multiple parts of your stack can be helpful, too.
- cryptonector 5y ago
- d--b 5y agoBack when I was in school in 2004, I had a teacher who had worked on the ASN.1 spec. In 2004, XML was all the rage. People would create "XML startups", and Microsoft did SOAP and some other guys XHTML, and XML schemas, semantic web and so on. I remember that teacher being so upset that XML got big and ASN.1 disappeared. It was very awkward. Poor guy...
- AceJohnny2 5y agoOur (computing) history is littered with better technology that was overtaken by "worst". I wonder if your teacher eventually understood why XML was preferred over ASN1. Seems to me like it was easier to pick up, and harder to mess up.
- cryptonector 5y agoTwo very funny things happened then: a) ASN.1 got XML Encoding Rules (XER), so you can use XML w/ ASN.1 as the schema language, which really, mostly is about supporting existing ASN.1-based protocols but with XML because well, you know, XML was all the rage, and b), FastInfoSet happened, which is an ASN.1 PER-based "compression" of XML because well, you know, XML is too verbose and unwieldy. I [bleep] you not, that happened. Evidence that there's nothing wrong with ASN.1 the syntax (and that's all it is, syntax and semantics, with a side of pluggable encoding rules where you can make them all up the way you want). Everything that's wrong with ASN.1 is either that which is wrong with BER/DER/CER (plenty), or that which is wrong with people's perception of ASN.1 (also plenty).
- dboreham 5y agoProto proto buf.
- lmilcin 5y agoI have worked for a time with credit card terminal applications. We used BER-TLV throughout the system extensively, where it was needed as well as where it wasn't. I have implemented complete parsers/serializers, data structures using TLV, transactional database where data was stored as TLV documents. EMV is built on top of BER-TLV, SSL used it, as well as ISO-8583 messages transmitted data encoded with BER-TLV. Communication with the PIN Pad was built on it. We kept configuration as BER-TLV documents. I have been able to parse hex representation in my head. I really liked the standard. It is nice, flexible and very efficient. Easy to parse, can be parsed reliably and safely in statically allocated memory. To those who think this is ancient history and it should be dropped -- do you think that might just be because you don't actually know it or maybe you just think it is old and so it must be bad?
- pitched 5y agoWhere EMV uses tags more like classes than types, I’m not really sure it actually counts as “abstract” syntax notation any more? Because all tags are these custom things, some don’t strictly parse out to unique type codes too. So a non-EMV parser will have a few tags that map to the same integer code and cause some fun bugs. That project was when I really understood deep-down why JSON won in the end!
- Woodi 5y ago> Can the veterans of the 90s SSL Wars explain the issues with ASN1/DER/BER? Looking it up today, it seems like a pretty smart and extensive serialization system, and I have to wonder why new systems like Google Protobufs chose to reinvent the wheel. Not a SSL or 90s veteran etc but: - ASN.1 is how OSI was try to jump on object orientation bandwagon - inheritance via text files declarations + types OIDs registration requirement via single entity somewhere on Earth... - ASN.1 is part of OSI - scientifically-correct attempt of networking standarisation - ASN.1 is part of OSI which was miraculously dropped exactly when Cold War ended and replaced by much simpler TCP/IP and friends. modulo security parts - that still need to use INSANE formats for passing numbers and strings... - ASN.1 implementations are guaranteed to be bugged for decades, imo and observations
- deleted 5y ago[deleted]
- cryptonector 5y agoI maintain Heimdal[0]'s ASN.1 compiler[1], though I didn't create it. It's a pleasure. It, and the IETF, have taught me a few things: - there's nothing really wrong with ASN.1 as a syntax except maybe it's ugly - there's nothing wrong at all with ASN.1's semantics - there's a TON wrong with the BER family of encoding rules (BER, DER, and CER), and with every tag-length-value scheme - you can create ASN.1 encoding rules for anything you like, which really means "use ASN.1 as the schema language for whatever encoding I prefer" - indeed, there's XER (XML encoding rules), JER (JSON encoding rules), GSER (generic string encoding rules) -- all text-based -- and a bunch of binary encodings with at least two that are not tag-length-value (and so resemble NDR and XDR), like PER and OER - people love to hate ASN.1, mainly because BER/DER/CER deserve the hatred, and for less legitimate reasons too, so they go off and invent new wheels that often have the same problems -- oh well! [0] https://github.com/heimdal/heimdal [1] https://github.com/heimdal/heimdal/tree/master/lib/asn1
- scoopr 5y agoIn the asn1 readme, and in some comments in these threads you mention the perils of the tag-length-value scheme, but you never seemed to explain whats wrong with it? At least in file formats it to me seem they would be instrumental to have a extendible and flexible format, where you can skip unknown or uninteresting chunks (in say, PNG chunks, or IFF-based formats like OBJ, etc.). Do you feel that the same doesn't apply to serialisation formats? How are the non-tlv binaries encoded then? Just implied offsets according to the schema? Can you then evolve the schema at all, or do you feel that both producer and consumer should have always access to the full schema, and flexiblity here is a non-feature? Sorry about the wall of questions, but I'm just so confused.
- memling 5y ago> In the asn1 readme, and in some comments in these threads you mention the perils of the tag-length-value scheme, but you never seemed to explain whats wrong with it? Not OP, but one of the challenges is that definite-length encodings like DER have to be encoded in a non-intuitive way. Values must be encoded prior to lengths (because the length is unknown), and the values can be nested. Therefore you have to encode a message essentially backwards when using definite-length encodings. This can potentially require a great deal of memory and can increase latency because streaming the data is hard. Indefinite lengths (BER has this option, CER requires it) can help avoid this problem, but then you lose the benefit of skipping elements (which you allude to in your next paragraph). > Do you feel that the same doesn't apply to serialisation formats? How are the non-tlv binaries encoded then? Just implied offsets according to the schema? Can you then evolve the schema at all, or do you feel that both producer and consumer should have always access to the full schema, and flexiblity here is a non-feature? You've hit the tradeoffs pretty well in the question, I think. The nice thing about TLV is that you can decode without a schema and potentially work with the contents: it's a relatively simple format to decode and validate even if it's not necessarily great for the encoder. ASN.1 supports schema-informed packed encodings that place greater demands on both the encoder and decoder. The main advantage is that they greatly reduce message overhead, but it requires a lot of bit-twiddling for presence/absence, default values, and, in unaligned variants, everything else, too. It's impossible, generally, to decode everything without the schema. PER has rules that disambiguate the values (e.g., they have to be ordered in a particular way, so you know what's coming next), and this mitigates some of the problems of TLV-style encodings. The tradeoffs are worth it when your pipes are small. 3GPP and LTE messages are largely encoded in PER. The people playing in that world usually have plenty of money to spend on commercial solutions and have bandwidth to roll their own, too. That's a bit different than smaller shops who are looking for convenient automated serialization formats.