3 ms·
I agree with Your claim about ASN.1 not intended to be parsed using hand written code. The ASN.1 is indeed quite close to a grammar specification, as also shown
by cryptonick 8y ago
I agree with Your claim about ASN.1 not intended to be parsed using hand written code. The ASN.1 is indeed quite close to a grammar specification, as also shown in the paper. However, I believe that a major issue related to its parsing complexity is the binary encoding generally used, which is either BER or DER: both of them employing length fields. While the usage of length fields is usually ubiquitous in formats related to communications protocols, these fields are quite annoying to be handled from a grammar design perspective. Indeed, a length field requires to count bytes of the payload: this operation is tedious to be done with grammars, while it is extremely easy with hand written code, in turn generally making the usage of grammar based automatic parser generators for these formats a less common choice.
From a grammar design standpoint, a delimiter based structure would be preferable. For instance, in the context of X.509, we proposed a new format[1] which replaces DER encoding with a new format where there are no longer length fields but the payload is terminated by a fixed delimiter. The grammar for this format was way easier than the length field base encoding, without requiring any hand written code.
[1] A novel regular fomat for X.509 digital certificates, https://link.springer.com/chapter/10.1007/978-3-319-54978-1_18 https://link.springer.com/chapter/10.1007/978-3-319-54978-1_...
- wahern 8y agoI believe I read that paper recently :) Ultimately I ended up using PEGs. LPeg in particular, using LPeg's match time captures to recursively invoke PEGs for length-encoded objects. (In addition to the match time capture extension, what's especially nice about LPeg--missing from every other PEG library I've seen--is that you can build and transform the AST in one shot.) I've also tentatively rejected translation to a format like that proposed in the paper. In a secure enclave-like environment I'd rather be dealing with statically defined C-like structs with stronger invariants--i.e. no optional or sum types, no variable-length fields; basically, no need for any kind of parsing whatsoever. If I have to transform, I'd like to transform the message both syntactically and semantically into the simplest possible form. Parsing complexity is only part of the equation. The other part is semantic complexity, which is a different kind of problem that better formats and parsers can't fix. In another timeline things could have been different, but we don't live in that timeline :( We can't let perfect be the enemy of good. Even if we could move away from DER or even ASN.1 in the open source world, the entire telecommunications industry (and specifically the cell industry) is built around ASN.1 and DER/PER/XER. AFAICT the biggest users of asn1c are people working with 3GPP and similar standards. No matter how sane and secure we can make our open source ecosystems, ASN.1 and similar older tech will still lurk in the background, remaining the weakest link in the chain. If we want real security we have no choice but to develop better tooling in that regard. I appreciate your proposal is very much of that mindset, I'm just not sold on the practical utility.