3 ms·
But you won't know at parse-time how to tread the item. Below we have duplicate "keys" and an atom in the middle of a k/v list -- what do we do here? Are the du
by thelazydogsback 6y ago
But you won't know at parse-time how to tread the item.
Below we have duplicate "keys" and an atom in the middle of a k/v list -- what do we do here? Are the dup keys an error, or do we make a multi-map? Do we throw away the atom or assume it's a key with a nil value?
((k1 v1)(k1 v2) ... (k4532 v4532) anAtom (k4533 v4533) ...)
It's stuff like this that makes Erlang/Elixir really weird -- it needs to look at the format of every item first to determine what it's going to be or how it's going to be printed. (Is it bunch of numbers, or a string? Is it a dictionary, or just a list of keys and values?)
- chriswarbo 6y ago> But you won't know at parse-time how to tread[sic] the item. I would follow the advice from LangSec ( http://langsec.org http://langsec.org ) and "parse, don't validate" ( https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va... ): - Parsing isn't just turning bytes into some generic tree representation (JSON/SAX/s-expr/etc.); it also includes subsequent domain/application specific input handling, e.g. constructing custom objects, checking certain invariants, etc. Note that we don't need custom objects; but we should still be checking the required invariants of our 'List[Tuple[String, Bool]]', or whatever. - Parsing should be done up-front, processing should use its output; e.g. we shouldn't be passing around generic types like 'JsonObject', 'AssocList', etc. (unless our application is generic JSON/s-expr processor, of course!). Note that we can still do streaming/lazy processing, with data parsed on-demand, but we should be careful about what side-effects might be performed part-way-through a broken input (this would apply to any approach though, e.g. hitting invalid bytes part way through some JSON) Examples like your broken assoc-list will hence be spotted at parse time; maybe not at the bytes->tree step, but certainly at the tree->domain-model step.
- thelazydogsback 6y agoI agree that that is how a parsing & instantiating framework should be built, and you should be able to hook in at any level -- but the reality is that normally when one consumes such data you are using both functions at the same time - de-serializing. Having list vs. set vs map notations is both syntax (for parsing) and really a form of meta-data, making explicit what promises are made of the data, and therefore hinting at which data-structs you may want to instantiate from the data. I suppose that ideally the meta-data should be a separate, optional statement made about the data, but that's not the M.O. for any common formats.