19 ms·
Parse, don't validate (2019)
- kellysutton 4y agoThis post resonates with a lesson I’ve learned in my career so far: It is always easier to relax constraints than tighten them.
- aeonik 4y agoWhat you say makes theoretical sense, but many bank systems still enforce weak password constraints because someone enforced those weak constraints 30 years ago in mainframe code that nobody seems to want to update.
- bruce343434 4y agoNote that this basically requires your language to have ergonomic support for sum types, immutable "data classes", pattern matching. The point is to parse the input into a structure which always upholds the predicates you care about so you don't end up continuously defensively programming in ifs and asserts.
- mtlynch 4y agoI get a lot of value from this rule even without those language features. I follow "Parse, Don't Validate" consistently in Go. For example, if I need to parse a JSON payload from an end-user for Foo, I define a struct called FooRequest, and I have exactly one function that creates a FooRequest instance, given a JSON stream. Anywhere else in my application, if I have a FooRequest instance, I know that it's validated and well-formed because it had to have come from my FooRequest parsing function. I don't need sum types or any special language features beyond typing.
- jotaen 4y agoMy main take-away is the same, I wonder though whether “parse, don’t validate” is the right term for it. To me, “parse, don’t validate” somehow suggests that you should do parsing instead of validation, but the real point for me is that I still validate (as before), plus I “capture”/preserve validation success by means of a type.
- qsort 4y agoIt's in the same sense of "whitelist, don't blacklist", or "by the love of god it's 2023, do not escape SQL". Don't define reasons why the input is invalid, instead have a target struct/object, and parse the input into that object.
- blincoln 4y agoI like this explanation and approach, but how does it solve the first problem described in the article - the case where there's an array being processed that might be empty? There are plenty of cases in real-world code where an array that's part of a struct or object may or may not contain any elements. If you're just parsing input into that, it seems like you'd either still end up doing an equivalent of checking whether the array is empty or not everywhere the array might be used later, even if that check is looking at an "array has elements" type flag in the struct/object, and so you're still maintaining a description of ways that the input may be invalid. But I'm not a world-class programmer, so maybe I'm missing something. Maybe you mean something like for branches of the code that require a non-empty array, you have a second struct/object and parser that's more strict and errors out if the array is empty?
- strgcmc 4y agoRemember, the author of the article constructed a scenario where, it was expected that the "main" function ends up treating an empty "CONFIG_DIRS" input as an uncatchable IOError; in other words, an empty array was invalid/not-allowed, per the rules of this program. Depending on the context in which you are operating, you may or may not have similar rules or requirements to follow. Empty lists are actually generally not a big deal - they are just lists of size 0, and they generally follow all the same things you can do with non-empty lists. The fact that a "head" function throws an error on an empty list, is really just a specific form of the more general observation that: any array would throw an index-out-of-bounds exception when given an index that's... out of of bounds. So any time you are dealing with arrays, you probably need to think about, "what happens if I try to index something that's out of bounds? is that possible?" In this particular contrived example, all that mattered was the head of the array. But what if you wanted to pick out the 3rd argument in a list of command line arguments, and instead the user only gave you 2 inputs? If 3 arguments are required, then throw an IOError as early as possible after failing to parse out 3 arguments; but once you pass the point of parsing the input into a valid object/struct/whatever, from that point forward you no longer care about checking whether the 3rd input is empty or not. So again, it depends on your scenario. Actually the more interesting variant of this issue (in OO languages at least) is probably handling nulls, as empty lists are valid lists, but nulls are not lists, and requires some different logic usually (and hence why NullPointerExceptions aka NPEs are such a common failure mode).
- jim-jim-jim 4y agoI think we'll eventually come to regard `if` as we do `goto`.
- leetrout 4y agoUsing pattern matching instead or something else?
- Joeri 4y agoBy using reactive programming techniques the program can be approached as a set of data streams mapping input to output, and conditional behavior becomes the application of different filters and combiners on the streams. This dovetails nicely with functional programming, which allows generic expression and reuse of those stream operations.
- oslac 4y agoNot a perfect example, but this can be seen (pattern match replacing if) with Kotlin's when.
- lmm 4y agoMost of the time you avoid having booleans in the first case, in favour of polymorphism (e.g. rather than having an "addOrMultiply" flag, you have separate "Add" and "Multiply" classes with a polymorphic method that does the addition or multiplication). You probably need some conditional logic in your "parser" (and whether that's "if" or pattern matching isn't so important IMO), but you should push booleans out of your core business logic and over to the edges of your program.
- leetrout 4y agoThat sounds miserable. Is there blog post or something with more details that supports this? I might be having a knee jerk reaction because I can't imagine something like this being easy to work with and maintain but I recognize you were just giving a trivial example.
- ocharles 4y agoThis isn't strictly true, an alternative is to have a language with enough encapsulation that you can parse into something that can only be observed as correct. The underlying parsing doesn't have to parse into sum types, provided your observation functions always preserve the parsed invariants.
- crabbone 4y agoIt's not just about these limitations. In order to be useful, type systems need to be simple, but there's no such restrictions on rules that govern our expectations of data correctness. OP is delusional if they think that their approach can be made practical. I mean, what if the expectation from the data that an value is a prime number? -- How are they going to encode this in their type systems? And this is just a trivial example. There are plenty of useful constraints we routinely expect in message exchanges that aren't possible to implement using even very elaborate type systems. For example, if we want to ensure that all ids in XML nodes are unique. Or that the last digit of SSN is a checksum of the previous digits using some complex formula. I mean, every Web developer worth their salt knows that regular expressions are a bad idea for testing email addresses (which would be an example of parsing), and it's really preferable to validate emails by calling a number of predicates on them. And, of course, these aren't the only examples: password validation (the annoying part that asks for capital letter, digit, special character? -- I want to see the author implement a parser to parse possible inputs to password field, while also giving helpful error messages s.a. "you forgot to use a digit"). Even though I don't doubt it's possible to do that, the resulting code would be an abomination compared to the code that does the usual stuff, i.e. just checks if a character is in a set of characters.
- flupe 4y agoBoth your examples (is my number prime, are my XML nodes unique) are easily expressed in a dependently-typed language. Dependent type checkers may be hard to implement, but the typing rules are fairly simple, and people have been using this correct by construction philosophy using dependently-typed languages for a while now. There's nothing delusional about that.
- crabbone 4y ago> dependently-typed language. Now, are there really tools to make type systems with dependent types simple to prove? In reasonable time? How about the effort developers would have to put into just trying to express such types and into verifying that such an expression is indeed accomplishing its stated goals? Just for a moment, imagine you filing a PR in a typical Web shop for the login form validation procedure, and sending a couple of screenfulls of Coq code or similar as a proof of your password validation procedure. How do you think your team will react to this? Again, I didn't say it's impossible. Quite the opposite, I said that it is possible, but will be so bad in practice that nobody will want to use it, unless they are batshit crazy.
- Thaxll 4y agoYou're describing deserialization in a strong typing language, sometimes it's not enough, ok your email went to an empty string which is useless.
- Kinrany 4y agoYou can have a ValidEmail type that performs all the checks on construction.
- Thaxll 4y agoOP said to not validate.
- lkitching 4y agoThe OP is contrasting between a 'validation' function with type e.g. validateEmail :: String -> IO () and a 'parsing' function validateEmail :: String -> Either EmailError ValidEmail The property encoded by the ValidEmail type is available throughout the rest of the program, which is not the case if you only validate.
- mrkeen 4y agoThat's fine. Validation would be: Email email = new Email(anyString); email.validate(); Parsing (in OP's context) would be: Either<Error, ValidEmail> eEmail = Email.from(anyString);
- gnulinux 4y agoYou're misunderstanding. Validation looks like validateEmail : String -> String -- post-condition: String contains valid email whereas parse looks like: parseEmail : String -> Either EmailError ValidEmail There is no problem using `ValidEmail` abstraction. The problem is type stability, when your program enters a stronger state at runtime (i.e. certain validations are performed at runtime) it's best to enter a strong state at compile time (stronger types) so that compiler can verify these conditions. If you remain at String, these validations (that a string is valid email) have no compile-time counterpart so there is no way for compiler to verify. So use `ValidEmail` instead.
- PartiallyTyped 4y agoI don't think sum types is necessary. TypeScript's gradual typing should suffice to capture this.
- PartiallyTyped 4y agoI was wrong. See this fantastic reply by the author: https://news.ycombinator.com/reply?id=35059519 https://news.ycombinator.com/reply?id=35059519
- jameshart 4y agoThose aren’t exactly rare features these days though. Beyond the functional world, TypeScript, C# and Java all have them to some extent, so it’s basically conquered all the mainstream object oriented languages. There’s even a proposal to add pattern matching to C++.
- armchairhacker 4y agoI wish more languages had some equivalent of records, tagged unions, and pattern matching. Don't have to be 100% immutable or perfect ADTs: see Rust, Swift, Kotlin. Even TypeScript can do this, albeit it's uglier with untagged unions and flow typing instead of pattern matching.
- swsieber 4y ago> The point is to parse the input into a structure which always upholds the predicates you care about so you don't end up continuously defensively programming in ifs and asserts. While the article is titled "parse don't validate" I like it's first point of make illegal states unrepresentable much better.
- deleted 4y ago[deleted]
- lolinder 4y agoSum types are used here for error handling, but if your language has a different error handling convention you can and should just use that. In Java, you'd implement this by making a class with a private constructor, no mutator methods, and a static factory method that throws an exception if the parsing fails. Since the only way to get an instance of the class is through the factory method, you've made illegal states unrepresentable and know that the class always holds to its invariants. No methods on instances of that class will throw exceptions from then on, so you've successfully applied "Parse, Don't Validate" without needing sum types. The point of the article isn't the particular implementation in Haskell, it's the concept of pushing all data error states to the boundaries of your code, which applies anywhere as long as you translate it into the idioms of your language.
- lexi-lambda 4y ago> In Java, you'd implement this by making a class with a private constructor, no mutator methods, and a static factory method that throws an exception if the parsing fails. This is similar, and is indeed quite useful in many cases, but it’s not quite the same. I explained why in this comment: https://news.ycombinator.com/item?id=35059886 https://news.ycombinator.com/item?id=35059886 (The comment is talking about TypeScript, but really everything there also applies to Java.)
- lolinder 4y agoThanks for the reply! I wasn't at all expecting one from you. If I'm understanding the difference correctly, it's that the constructive data modeling approach can be proven entirely in the type system without any trust in the library code, while the Java approach I recommended depends on there being no other way to construct an instance of the class, which can be tricky to guarantee. Is that accurate?
- lexi-lambda 4y agoYes, that’s about right. But really do read the followup blog post (https://lexi-lambda.github.io/blog/2020/11/01/names-are-not-type-safety/ https://lexi-lambda.github.io/blog/2020/11/01/names-are-not-...), as it explains that in much more depth! In particular, it says: > To some readers, these pitfalls may seem obvious, but safety holes of this sort are remarkably common in practice. This is especially true for datatypes with more sophisticated invariants, as it may not be easy to determine whether the invariants are actually upheld by the module’s implementation. Proper use of this technique demands caution and care: > * All invariants must be made clear to maintainers of the trusted module. For simple types, such as NonEmpty, the invariant is self-evident, but for more sophisticated types, comments are not optional. > * Every change to the trusted module must be carefully audited to ensure it does not somehow weaken the desired invariants. > * Discipline is needed to resist the temptation to add unsafe trapdoors that allow compromising the invariants if used incorrectly. > * Periodic refactoring may be needed to ensure the trusted surface area remains small. It is all too easy for the responsibility of the trusted module to accumulate over time, dramatically increasing the likelihood of some subtle interaction causing an invariant violation. > In contrast, datatypes that are correct by construction suffer none of these problems. The invariant cannot be violated without changing the datatype definition itself, which has rippling effects throughout the rest of the program to make the consequences immediately clear. Discipline on the part of the programmer is unnecessary, as the typechecker enforces the invariants automatically. There is no “trusted code” for such datatypes, since all parts of the program are equally beholden to the datatype-mandated constraints. They are both quite useful techniques, but it’s important to understand what you’re getting (and, perhaps more importantly, what you’re not).
- philsnow 4y ago> this basically requires your language to have ergonomic support for sum types, immutable "data classes", pattern matching Not at all, the article is about pushing complexity to the "edges" of your code so that the gooey center doesn't have to faff around with (re-)checking the same invariants over and over... but its examples are also in Haskell, in which it would be weird to do this without the type system. In python or java or whatever you'd just parse your received_api_request_could_be_sketchy or fetched_db_records_but_are_they_really into a BlessedApiRequest or DefinitelyForRealDBRecords in their constructors or builder methods or whatever, disallow any other ways of creating those types, and then exclusively using those types. edit: wait, actually no we agree, I must have glossed over your second sentence, sorry
- not_knuth 4y agoPreviously on HN: 2 years ago: https://news.ycombinator.com/item?id=27639890 https://news.ycombinator.com/item?id=27639890 3 years ago: https://news.ycombinator.com/item?id=21476261 https://news.ycombinator.com/item?id=21476261
- Pulz 4y agoIt's at least once per month.
- dang 4y agoThanks! Macroexpanded: Parse, Don't Validate (2019) - https://news.ycombinator.com/item?id=27639890 https://news.ycombinator.com/item?id=27639890 - June 2021 (270 comments) Parse, Don’t Validate - https://news.ycombinator.com/item?id=21476261 https://news.ycombinator.com/item?id=21476261 - Nov 2019 (230 comments) Parse, Don't Validate - https://news.ycombinator.com/item?id=21471753 https://news.ycombinator.com/item?id=21471753 - Nov 2019 (4 comments)
- artagnon 4y agoThe example described in this post is JSON parsing in Haskell, but I've implemented a complicated compiler transform that lifts loops to static control parts (SCOPs), in the past, in C++. Each inner function in the lift would switch on valid constructs, either returning a lifted integer set, or throwing an exception on match failure. Although exceptions have a non-trivial cost in C++, it was the cleanest design I could come up with at the time.
- tizzy 4y agoI always favourite this when it comes up on HackerNews, a great principle to follow or at least have in mind all the time
- conaclos 4y agoWhile theoretically I accepted with all the points, in practice it is sometimes too bloated to add so many types that are barely distinct. It is sometimes better to trade safety for simplicity. The trade-off is always hard to make. For instance: should I introduce a branded type for unsigned 32bit integer in TypeScript? type u32 = number & { [U32_BRAND]: never } const u32 = (n: number): u32 => (n >>> 0) as u32 And then to make it hard any use of this type: declare let x, y: u32 y = u32(x + 1)
- Hackbraten 4y agoI just learned that there’s an open issue [0], apparently for introducing a similar feature. [0]: https://github.com/microsoft/TypeScript/issues/43505 https://github.com/microsoft/TypeScript/issues/43505
- oslac 4y agoOne of the best articles about programming tbh.
- conaclos 4y agoThe same author wrote a follow-up article [0] "Names are not type safety". [0] https://lexi-lambda.github.io/blog/2020/11/01/names-are-not-type-safety/ https://lexi-lambda.github.io/blog/2020/11/01/names-are-not-...
- ckdot2 4y agoPlease, don't write your own JSON parser/validator. There's JSON Schema https://json-schema.org https://json-schema.org which has implementations in most languages. You can valiate your JSON by a given, standardized JSON schema file - and you're basically done. After the validation, it's probably good practise to map the JSON to some DTO and may do some further validation which doesn't check the structure of the data but it's meaning.
- mirekrusin 4y agoJson schema doesn't have relation with static type system, ie. in typescript it's much better to use composable, functional combinators at i/o boundaries only and don't do any extra checks anywhere where type system provides guarantees.
- ckdot2 4y agoI think it's good enough. Besides JSON Schema being a standard instead of custom solution, you also get nice error messages in case there's a validation issue. If your JSON schema file is properly defined it should be safe enough to just map your JSON into some static type DTO afterwards and trust your data and it's types to be valid. In JSON Schema you can validate for strings, numbers, integers, and custom objects. It's quite powerful and - personally - I wouldn't want to implement that kind of stuff on my own.
- mirekrusin 4y agoYou don't need to implement it on your own, you can use library. Nice error messages exist there as well. If you're casting untyped results, you can change one side and not the other and find out about this problem when in production. Or simply any mistake will get unnoticed. Using typescript first library allows you to do much more - supports opaque types, custom constructors and any imaginable validation that can't be expressed in json schema.
- lexi-lambda 4y agoAmusingly, the tweet that inspired this blog post—which is linked in the second paragraph of the article—is specifically about how automatically generating a JSON parser from your datatypes means you don’t have to implement that kind of stuff on your own, and there is no possibility of some separate “schema” going out of sync with your application logic. Of course, if you want to share the schema with downstream clients so that other programs can use it, that is a great use case for something like JSON Schema. It is a common interface that allows two different programs—quite possibly written in completely different languages—to communicate using the same format. That’s great! But it’s only half the story, because just having the schema doesn’t help you in any way to make sure the code actually respects that schema. That’s where integration with the language’s type system can help, perhaps by automatically generating types from the schema and then generating parsing/serialization functions that use those generated types.
- agumonkey 4y agoI don't know who else sees everything as language problems. Even REST seems like a concrete syntax over HTTP, mostly relational, resources
- bob1029 4y agoThis is something I've become fairly passionate about lately. Any time I see some regex, I start asking probing questions about the nature of the underlying abstraction. Being able to deterministically convert something into an AST is the ultimate test of that thing's stability at any scale.
- thanatropism 4y agoThat's also kind of an IQ test.
- arminsergiony 4y agoI believe that JSON Schema is a great solution because it is a standardized format and provides helpful error messages for validation issues. If the schema file is well-defined, it should be safe to map the JSON data to a static type DTO and trust that the data types are valid. JSON Schema's ability to validate strings, numbers, integers, and custom objects makes it a powerful tool, and I personally wouldn't want to attempt to implement something similar on my own.
- redbar0n 4y agoSeveral earlier HN threads about this article: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=%22parse%2C%20don%27t%20validate%22&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- jameshart 4y agoThis post absolutely captures an essential truth of good programming. Unfortunately, it conceals it behind some examples that - while they do a good job of illustrating the generality of its applicability - don’t show as well how to use this in your own code. Most developers are not writing their own head method on the list primitive - they are trying to build a type that encapsulates a meaningful entity in their own domain. And, let’s be honest, most developers are also not using Haskell. As a result I have not found this article a good one to share with junior developers to help them understand how to design types to capture the notion of validity, and to replace validation with narrowing type conversions (which amount to ‘parsing’ when the original type is something very loose like a string, a JSON blob, or a dictionary). Even though absolutely those practices follow from what is described here. Does anyone know of a good resource that better anchors these concepts in practical examples?
- asimpletune 4y agoI think it's really hard to learn from reading unfortunately. It's one of those things where if you get it, you get it, but it kind of takes personal experience to fully grok it. I guess because there are a lot of subtle differences.
- epolanski 4y ago> And, let’s be honest, most developers are also not using Haskell. Everything in that post applies to the most common programming language out there: TypeScript. And several popular others such as Rust, Kotlin or Scala.
- elfprince13 4y agoIt also applies to C++ and Java!
- tialaramex 4y agoAnd parse-don't-validate is often very nice to work with in Rust, I can describe how to turn some UTF-8 text into my type Foo in a function: impl std::str::FromStr for Foo { type Err = ReasonsItIsNotAFoo; fn from_str(s: &str) -> Result<Self, Self::Err> { /* etc. */ } } And then whenever I've got a string which I know ought to be a Foo, I can: let foo: Foo = string.parse().expect("This {string:?} ought to be a Foo but it isn't"); Since we said foo is a Foo, by inference the parsing of string needs to either succeed with a Foo, or fail while trying, so it calls that FromStr implementation we wrote earlier to achieve that.
- roenxi 4y agoparseNonEmpty [] = throwIO $ userError "list cannot be empty" How would that interact with a scenario where we want a specific error message if a specific list is empty? Eg, "you want to build the list using listBuilder()". Making illegal states unrepresentable is good advice but I don't think that escapes the value of good validation. It is a mistake to do ad-hoc validation. But it makes a lot of sense to have a validation phase, a parse phase then an execution phase when dealing with untrusted data. The validation phase gives context-aware feedback, the parse phase catches what is left and then execution happens. A type system doesn't seem like a good defence against end user error. The error messages in practice are mystic. I think the complaint here is if people are trying to implement a type system using ad-hoc validation which is a bad idea.
- jameshart 4y agoWhen you’re building a ‘parser’ (in this broad, type-narrowing sense) to handle user-supplied data, the result type really needs to be a rich mix of - successfully parsed data objects - error objects - warning objects That way your consumers can themselves decide what to do in the face of errors and warnings. (Of course one ugly old fashioned way to add optional ‘error’ types to your return signature is checked exceptions, but we don’t talk about that model any more.)
- lexi-lambda 4y agoCertainly I don’t think `parseNonEmpty` would be especially useful in a real program, it’s only there as an example to provide a particularly simple contrasting example against `validateNonEmpty`. The example earlier in the blog post using the `nonEmpty` function (which returns an optional result) is a more realistic example of how such things are actually used in practice, since that allows you to raise a domain-appropriate error message. Tangentially, in Haskell specifically, I have actually written a library specifically designed for checking the structure of input data and raising useful error messages, which is somewhat ironically named `monad-validate` (https://hackage.haskell.org/package/monad-validate https://hackage.haskell.org/package/monad-validate). But it has that name because similar types have historically been named `Validation` within the Haskell community; using the library properly involves doing “parsing” in the way this blog post advocates.
- kybernetikos 4y agoThis is obviously good advice almost all of the time. However, I have had to deal occasionally with http libraries that tried to parse everything and would not give you access to anything that they could not parse. This was incredibly frustrating for corner cases that the library authors hadn't considered. If you are the one who is going to take action on the data, parse don't validate is the correct approach. If you are writing a library that deals with data that it doesn't fully understand, and you're handing that data to someone else to take action with, then it may not always be the right approach.
- tizzy 4y agoThis seems like good library design. As annoying as it is, it means the things you can use are well tested and supported. What was your solution to this? Parse the things the library didn't?
- kybernetikos 4y agoThe library didn't allow you to see the things (e.g. particular headers or options for those headers) that it didn't know to parse. Ultimately we had to migrate to a different library that didn't restrict us to just what the library knew. The decision not to let us even see things that the library didn't know about is particularly egregious where best practices are changing over time. In my view it's a very bad design for an http library, although it would have been a lot less frustrating if it had at least provided an escape hatch.
- jimbokun 4y agoSounds like its model is not the HTTP RFC, but something more specific to some domain. Which I agree, is a poor design choice. A type modeling an HTTP request should model the RFC definition as closely as possible.
- aidenn0 4y ago> Which I agree, is a poor design choice. A type modeling an HTTP request should model the RFC definition as closely as possible. I couldn't disagree more. A type modeling an HTTP request should model HTTP requests. Not some theoretical description of an HTTP request.
- jmull 4y agoI get the point, but I wonder at why people find this particular article compelling. To me it's weak... It's built on a particular technical distinction between paring and validating that (1) is not all that commonly understood or consistently accepted and (2) not actually explicitly stated in the article! (validation: check data assumptions, fail of not met; parse: check data assumptions, fail if not met, and on success return data as a new type reflecting the additional constraints of the data, which can therefore be checked at compile time. Notice parsing includes validation, which makes the title of the article quite poor.) That's important to know because the distinction is only meaningful in the context of certain language features, which may or may not apply. Also, this is not great general advice: > Push the burden of proof upward as far as possible, but no further For one, it's a mostly meaningless, since it really just says put the burden of proof in the right place. But it implies that upward is preferable. You really want to push it upward if it's a high-level concern, and downward if it's a low-level concern. E.g., suppose you're working on an app or service that accesses the database, so the database is lower-level. You'll want to push your database-specific type transformations closer to the code that accesses the database. Honestly, I find this whole thing kind of muddled. (Also, in my experience, the fundamental limit here isn't on validation strategies, but the human ability to break down a problem and logically organize the solution. You can just as easily end up with an unmaintainable mess of spaghetti types as with any other useful abstraction.
- jakelazaroff 4y ago> You really want to push it upward if it's a high-level concern, and downward if it's a low-level concern. E.g., suppose you're working on an app or service that accesses the database, so the database is lower-level. You'll want to push your database-specific type transformations closer to the code that accesses the database. IMO, database code is at exactly the same level of concern as network code or filesystem code. By “upward”, she means push parsing to the boundaries of your program — as close to the point of ingress as possible.
- jmull 4y agoThe db access is just an example. I used upward and downward working off the terminology of the article. But I can put it like this: For a given call or request, there's input, some work done with that input, and the result. (This is true, whether we're talking about a functional or imperative style.) Your code will have some structure that reflects the work to be done. You want to push your parsing toward the input if it's concerned with the input, and toward the result if it's concerned with the result. Whether you want to call the processing closer to the input "upward", or "earlier" or whatever, that's fine with me. If you call the processing closer to the input and closer to the result both "upward" then I think it's not a useful metaphor and you should choose a different one.
- Octokiddie 4y agoI like how the author boils the idea down into a simple comparison between two alternative approaches to a simple task: getting the first element of a list. Two alternatives are presented: parseNonEmpty and validateNonEmpty. From the article: > The difference lies entirely in the return type: validateNonEmpty always returns (), the type that contains no information, but parseNonEmpty returns NonEmpty a, a refinement of the input type that preserves the knowledge gained in the type system. Both of these functions check the same thing, but parseNonEmpty gives the caller access to the information it learned, while validateNonEmpty just throws it away. This might not seem like much of a distinction, but it has far-reaching implications downstream: > These two functions elegantly illustrate two different perspectives on the role of a static type system: validateNonEmpty obeys the typechecker well enough, but only parseNonEmpty takes full advantage of it. If you see why parseNonEmpty is preferable, you understand what I mean by the mantra “parse, don’t validate.” parseNonEmpty is better because after a caller gets a NonEmpty it never has to check the boundary condition of empty again. The first element will always be available, and this is enforced by the compiler. Not only that, but functions the caller later calls never need to worry about the boundary condition, either. The entire concern over the first element of an empty list (and handling the runtime errors that result from failure to meet the boundary condition) disappear as a developer concern.
- waynesonfire 4y agoDoes this scale? What if you have 10 other conditions?
- ElevenLathe 4y agoI think what you're getting at is that it seems ponderous to have types named things like NonEmptyListWhereTheThirdElementIsTheIntegerFourAndTheOtherElementsAreStringsOfLengthSixOrUnder and the answer is that you shouldn't do that, but instead name it something in the problem domain (of whatever the program is about) like WidgetDescription or whatever.
- tialaramex 4y ago
- harveywi 4y agoBased on Conal Elliott's new formulation of two dual parsing implementations [1] using inductive regular expressions and coinductive tries, where the former corresponds to symbolic differentiation and the latter corresponds to automatic differentation, the dual statment of "parse, don't validate" might be: "Do or do not. There is no trie." [1] http://conal.net/papers/language-derivatives/ http://conal.net/papers/language-derivatives/
- asimpletune 4y agoOne way I like to use to understand and explain what the author is talking about is IKEA furniture that can only go together one way, the "right" way. People are inevitably going to not look at the manual and just start fiddling, so you design the pieces themselves to reject the wrong combinations. I don't know if Ikea actually does this, but I just mean as a concept that's one way you can use to imagine it. There are so many examples of this in the wild, for important things, e.g. you can't use the washing machine unless the lid is actually closed all the way.
- dgunay 4y agoSometimes they do, but not always - my roommate once assembled an Ikea bookshelf with one of the panels on backwards.
- cratermoon 4y agoCurrently working on a project with some inexperienced developers where I was brought on to consult. Lots of shotgun parsing. Stringly-typed code everywhere.Even the dates and timestamps are pass around as strings, resulting in code to go back and forth between proper time types and strings all over the place.
- frankreyes 4y agoMany years ago I had to write a code transformation for a legacy programming language. It was orders of magnitude much easier to assume that the code was syntactically valid. And we could make that assumption because we had the legacy compiler being used to compile.
- deleted 4y ago[deleted]
- marcelr 4y agoYes, and this is not tied to statically typed languages. If anything this is simpler to do in dynamic languages, but the culture isn’t there in my experience.
- romankolpak 4y agothis generalizes more broadly in fact and applies to not just parsing and validating data. very often you want to reject all problematic states so you can conveniently code the happy path assuming all preconditions are met (inputs are correct, permissions are granted, etc) using meaningful data structures free from the messiness of the real world. it's often a matter of experience to get this nuance of programming. you just learn with time that it's very inconvenient to test for emptiness multiple levels deep in the callstack again and again and you go "why can't i just assume good data here?". and then you figure out a way to write the code so you can.
- Joel_Mckay 4y agoIn general, parsers that do not limit recursion-depth and order can be a problem. Marshalling the data for platform traversal is also very wise. A library like Xalan/xerces using XSLT is very powerful, or something lightweight like the JSON/BSON parser in libbson. Accordingly, one must assume the data is _always_ malformed, and assign a scoring system to the expected format at each stage of decoding. i.e. each service/function does a sanity check, then scores which data is critical, optional, and prohibited. This way your infrastructure handles the case when (not if) someone tries to put Coffee Grounds in your garbage disposal unit. =)
- psychoslave 4y ago>it’s only three words long: Parse, don’t validate. My own English parser is telling me it's actually four words, however. Your mileage may validate that differently. ;)