10 ms·
Parse, Don't Validate (2019)
- seanwilson 8mo agoMaybe I'm missing something and I'm glad this idea resonates, but it feels like sometime after Java got popular and dynamic languages got a lot of mindshare, a large chunk of the collective programming community forgot why strong static type checking was invented and are now having to rediscover this. In most strong statically typed languages, you wouldn't often pass strings and generic dictionaries around. You'd naturally gravitate towards parsing/transforming raw data into typed data structures that have guaranteed properties instead to avoid writing defensive code everywhere e.g. a Date object that would throw an exception in the constructor if the string given didn't validate as a date (Edit: Changed this from email because email validation is a can of worms as an example). So there, "parse, don't validate" is the norm and not a tip/idea that would need to gain traction.
- bcrosby95 8mo agoIn my experience that's pretty rare. Most people pass around string phone numbers instead of a phonenumber class. Java makes it a pain though, so most code ends up primitive obsessed. Other languages make it easier, but unless the language and company has a strong culture around this, they still usually end up primitive obsessed.
- vips7L 8mo agorecord PhoneNumber(String value) {} Huge pain.
- kleiba 8mo agoWhat have you gained?
- jalk 8mo agoAn explicit type
- dylan604 8mo agoObviously the pseudo code leaves to the imagination, but what benefits does this give you? Are you checking that it is 10-digits? Are you allowing for + symbols for the international codes?
- munk-a 8mo agoThat's going to be up to the business building the logic. Ideally those assumptions are clearly encoded in an easily readable manner but at the very least they should be captured somewhere code adjacent (even if it's just a comment and the block of logic to enforce those restraints).
- bjghknggkk 8mo agoHow to make a crap system that users will hate: Let some architecture astronaut decide what characters should be valid or not.
- bjghknggkk 8mo agoAnd parentheses. And spaces (that may, or may not, be trimmed). And all kind of unicode equivalent characters, that might have to be canonicalized. Why not treat it as a byte buffer anyway.
- JambalayaJimbo 8mo agoIf you are not checking that the phone number is 10 digits (or whatever the rules are for the phone number for your use case), it is absolutely pointless. But why would you not?
- jghn 8mo agoI would argue it's the other way around. If I take a string I believe to be a phone number and wrap it in a `PhoneNumber` type, and then later I try to pass it in as the wrong argument to a function like say I get order of name & phone number reversed, it'll complain. Whereas if both name & phone number are strings, it won't complain. That's what I see as the primary value to this sort of typing. Enforcing the invariants is a separate matter.
- deleted 8mo ago[deleted]
- jonathanlydall 8mo agoI’m very much a proponent of statically typed languages and primarily work in C#. We tried “typed” strings like this on a project once for business identifiers. Overall it worked in making sure that the wrong type of ID couldn’t accidentally be used in the wrong place, but the general consensus after moving on from the project was that the “juice was not worth the squeeze”. I don’t know if other languages make it easier, but in c# it felt like the language was mostly working against you. For example data needs to come in and out over an API and is in string form when it does, meaning you have to do manual conversions all the time. In c# I use named arguments most of the time, making it much harder to accidentally pass the wrong string into a method or constructor’s parameter.
- Akronymus 8mo agoIn f# you can use a single case discriminated union to get that behaviour fairly cheaply, and ergonomically. https://fsharpforfunandprofit.com/posts/designing-with-types-single-case-dus/ https://fsharpforfunandprofit.com/posts/designing-with-types...
- yakshaving_jgt 8mo agoIt's a design choice more than anything. Haskell's type safety is opt-in — the programmer has to actually choose to properly leverage the type system and design their program this way.
- pjerem 8mo ago> In most strong statically typed languages, you wouldn't often pass strings and generic dictionaries around. In 99% of the projects I worked on my professional life, anything that is coming from an human input is manipulated as a string and most of the time, it stays like this in all of the application layers (with more or less checks in the path). On your precise exemple, I can even say that I never saw something like an "Email object".
- Boxxed 8mo agoWell that's terrifying
- tracker1 8mo agoWhat's funny, is this is exactly one of the reasons I happen to like JavaScript... at its' core, the type coercion and falsy boolean rules work really well (imo) for ETL type work, where you're dealing with potentially untrusted data. How many times have you had to import a CSV with a bad record/row? It seems to happen all the time, why, because people use and manually manipulate data in spreadsheets. In the end, it's a big part of why I tend to reach for JS/TS first (Deno) for most scripts that are even a little complex to attempt in bash.
- jghn 8mo agoI've seen a mix between stringly typed apps and strongly typed apps. The strongly typed apps had an upfront cost but were much better to work with in the long run. Define types for things like names, email address, age, and the like. Convert the strings to the appropriate type on ingest, and then inside your system only use the correct types.
- deleted 8mo ago[deleted]
- rileymichael 8mo agothis is likely an ecosystem sort of thing. if your language gives you the tools to do so at no cost (memory/performance) then folks will naturally utilize those features and it will eventually become idiomatic code. kotlin value classes are exactly this and they are everywhere: https://kotlinlang.org/docs/inline-classes.html https://kotlinlang.org/docs/inline-classes.html
- wat10000 8mo agoI'm not sure, maybe a little bit. My own journey started with BASIC and then C-like languages in the 80s, dabbling in other languages along the way, doing some Python, and then transitioning to more statically typed modern languages in the past 10 years or so. C-like languages have this a little bit, in that you'll probably make a struct/class from whatever you're looking at and pass it around rather than a dictionary. But dates are probably just stored as untyped numbers with an implicit meaning, and optionals are a foreign concept (although implicit in pointers). Now, I know that this stuff has been around for decades, but it wasn't something I'd actually use until relatively recently. I suspect that's true of a lot of other people too. It's not that we forgot why strong static type checking was invented, it's that we never really knew, or just didn't have a language we could work in that had it.
- conartist6 8mo agoI think you're quite right that the idea of "parse don't validate" is (or can be) quite closely tied to OO-style programming. Essentially the article says that each data type should have a single location in code where it is constructed, which is a very class-based way of thinking. If your Java class only has a constructor and getters, then you're already home free. Also for the method to be efficient you need to be able to know where an object was constructed. Fortunately class instances already track this information.
- Archelaos 8mo agoStrong static type checking is helpful when implementing the methodology described in this article, but it is besides its focus. You still need to use the most restrictive type. For example, uint, instead of int, when you want to exclude negative values; a non-empty list type, if your list should not be empty; etc. When the type is more complex, specific contraints should be used. For a real live example: I designed a type for the occupation of a hotel booking application. The number of occupants of a room must be positiv and a child must be accompanied by at least one adult. My type Occupants has a constructor Occupants(int adults, int children) that varifies that condition on construction (and also some maximum values).
- imtringued 8mo agoUsing uint to exclude negative values is one of the most common mistakes, because underflow wrapping is the default instead of saturation. You subtract a big number from a small number and your number suddenly becomes extremely large. This is far worse than e.g. someone having traveled a negative distance.
- Archelaos 8mo agoIn C# I use the 'checked' keyword in this or similar cases, when it might be relevant: c = checked(a - b); Note that this does not violate the "Parse, Don't Validate" rule. This rule does not prevent you from doing stupid things with a "parsed" type. In other cases, I use its cousin unchecked on int values, when an overflow is okay, such as in calculating an int hash code.
- lelanthran 8mo ago> The number of occupants of a room must be positiv and a child must be accompanied by at least one adult. My type Occupants has a constructor Occupants(int adults, int children) that varifies that condition on construction (and also some maximum values). Or, you could do what I did when faced with a similar problem - I put in a PostgreSQL constraint. Now, no matter which application, now or in the future, attempts to store this invalid combination, it will fail to store it. Doing it in code is just asking for future errors when some other application inserts records into the same DB. Business constraints should go into the database.
- css_apologist 8mo agoThis is an idea that is not ON or OFF You can get ever so gradually stricter with your types which means that the operations you perform on on a narrow type is even more solid It is also 100% possible to do in dynamic languages, it's a cultural thing
- jackpirate 8mo ago> Edit: Changed this from email because email validation is a can of worms as an example Email honestly seems much more straightforward than dates... Sweden had a Feb 30 in 1712, and there's all sorts of date ranges that never existed in most countries (e.g. the American colonies skipped September 3-13 in 1752).
- flqn 8mo agoDates are unfortunate in that you can only really parse them reliably with a TZDB.
- legulere 8mo agoIt’s a ISO-standard to use Gregorian dates even for dates predating its invention. If you need to support anything else (I never had to in my Eurocentric work so far), you’ll need to model calendars, similar to how temporal did for JavaScript: https://tc39.es/proposal-temporal/docs/calendars.html https://tc39.es/proposal-temporal/docs/calendars.html
- brooke2k 8mo agothis is very much a nitpick, but I wouldn't call throwing an exception in the constructor a good use of static typing. sure, it's using a separate type, but the guarantees are enforced at runtime
- zanecodes 8mo agoGiven that the compiler can't enforce that users only enter valid data at compile time, the next best thing is enforcing that when they do enter invalid data, the program won't produce an `Email` object from it, and thus all `Email` objects and their contents can be assumed to be valid.
- mh2266 8mo agoThis is all pretty language-specific and I think people may end up talking past each other. Like, my preferred alternative is not "return an invalid Email object" but "return a sum type representing either an Email or an Error", because I like languages with sum types and pattern matching and all the cultural aspects those tend to imply. But if you are writing Python or Java, that might look like "throw an exception in the constructor". And that is still better than "return an Email that isn't actually an email".
- zanecodes 8mo agoAh yeah, I guess I assumed by the use of the term "contructor" that GP meant a language like Python or Java, and in some cases it can difficult to prevent misuse by making an unsafe constructor private and only providing a public safe contructor that returns a sum type. I definitely agree returning a sum type is ideal.
- imtringued 8mo agoI agree and for several reasons. If you have onerous validation on the constructor, you will run into extremely obvious problems during testing. You just want a jungle, but you also need the ape and the banana.
- noelwelsh 8mo agoIn 2 out of 3 problematic bugs I've had in the last two years or so were in statically typed languages where previous developers didn't use the type system effectively. One bug was in a system that had an Email type but didn't actually enforce the invariants of emails. The one that caused the problem was it didn't enforce case insensitive comparisons. Trivial to fix, but it was encased in layers of stuff that made tracking it down difficult. The other was a home grown ORM that used the same optional / maybe type to represent both "leave this column as the default" and "set this column to null". It should be obvious how this could go wrong. Easy to fix but it fucked up some production data. Both of these are failures to apply "parse, don't validate". The form didn't enforce the invariants it had supposedly parsed the data into. The latter didn't differentiate two different parsing.
- rzwitserloot 8mo agothat's a bit of a hairy situation. You're doing it wrong. Or not really, but.. complicated. As per [RFC 5321](https://www.rfc-editor.org/rfc/rfc5321.html https://www.rfc-editor.org/rfc/rfc5321.html): > the local-part MUST be interpreted and assigned semantics only by the host specified in the domain part of the address. You're not allowed to do that. The email address `foo@bar.com` is identical to `foo@BAR.com`, but not necessarily identical to `FOO@bar.com`. If we're going to talk about 'commonly applied normalisations at most email providers', where do you draw that line? Should `foo+whatever@bar.com` be considered equal to `foo@bar.com`? That souds weird, except - that is exactly how gmail works, a couple of other mail providers have taken up that particular torch, and if your aim is to uniquely identify a 'recipient', you can hardcode that `a@gmail.com` and `a+whatever@gmail.com` definitely, guaranteed, end up at the same mailbox. In practice, yes, users _expect_ that email addresses are case insensitive. Not just users, even - various intermediate systems apply the same incorrect logic. This gets to an intriguing aspect of hardcoding types: You lose the flex, mostly. types are still better - the alternative is that you reliably attempt to write the same logic (or at least a call to some logic) to disentangle this mess every time you do anything with a string you happen to know is an email address which is terrible but gives you the option of intentionally not doing that if you don't want to apply the usual logic. That's no way to program, and thus actual types and the general trend that comes with it (namely: We do this right, we write that once, and there is no flexibility left). Programming is too hard to leave room for exotic cases that programmers aren't going to think about when dealing with this concept. And if you do need to deal with it, it can still be encoded in the type, but that then makes visible things that in untyped systems are invisible (if my email type only has a '.compare(boolean caseSensitive)' style method, and is not itself inherently comparable because of the case sensitivity thing, that makes it _seem_ much more complicated than plain old strings. This is a lie - emails in strings *IS* complicated. They just are. You can't make that go away. But you can hide it, and shoving all data in overly generic data types (numbers and strings) tends to do that.
- masklinn 8mo ago> it feels like sometime after Java got popular [...] a large chunk of the collective programming community forgot why strong static type checking was invented and are now having to rediscover this. I think you have a very rose-tinted view of the past: while on the academic side static types were intended for proof on the industrial side it was for efficiency. C didn't get static types in order to prove your code was correct, and it's really not great at doing that, it got static types so you could account for memory and optimise it. Java didn't help either, when every type has to be a separate file the cost of individual types is humongous, even more so when every field then needs two methods. > In most strong statically typed languages, you wouldn't often pass strings and generic dictionaries around. In most strong statically typed languages you would not, but in most statically typed codebases you would. Just look at the Windows interfaces. In fact while Simonyi's original "apps hungarian" had dim echoes of static types that got completely washed out in system, which was used widely in C++, which is already a statically typed language.
- guerrilla 8mo ago> I think you have a very rose-tinted view of the past I think they also forgot the entire Perl era.
- esafak 8mo agoThat's understandable. Youthful indiscretion is best forgotten.
- zahlman 8mo agoI can still remember trying to deal with structured binary data in Perl, just because I didn't want to fiddle around with memory management in C. I'm not sure it was actually any less painful, and I ultimately abandoned that first attempt. (Decades later, my "magnum opus" has been through multiple mental redesigns and unsatisfactory partial implementations. This time, for sure...)
- chriswarbo 8mo ago> You'd naturally gravitate towards parsing/transforming raw data into typed data structures that have guaranteed properties instead to avoid writing defensive code everywhere e.g. a Date object that would throw an exception in the constructor if the string given didn't validate as a date It's tricky because `class` conflates a lot of semantically-distinct ideas. Some people might be making `Date` objects to avoid writing defensive code everywhere (since classes are types), but... Other people might be making `Date` objects so they can keep all their date-related code in one place (since classes are modules/namespaces, and in Java classes even correspond to files). Other people might be making `Date` objects so they can override the implementation (since classes are jump tables). Other people might be making `Date` objects so they can overload a method for different sorts of inputs (since classes are tags). I think the pragmatics of where code lives, and how the execution branches, probably have a larger impact on such decisions than safety concerns. After all, the most popular way to "avoid writing defensive code everywhere" is to.... write unsafe, brittle code :-(
- munificent 8mo ago> You'd naturally gravitate towards parsing/transforming raw data into typed data structures that have guaranteed properties instead to avoid writing defensive code everywhere e.g. There's nothing natural about this. It's not like we're born knowing good object-oriented design. It's a pattern that has to be learned, and the linked article is one of the well-known pieces that helped a lot of people understand this idea.
- thom 8mo agoMy experience was that enterprise programmers burned out on things like WSDL at about the same time Rails became usable (or Django if you’re that way inclined). Rails had an excellent story for validating models which formed the basis for everything that followed, even in languages with static types - ASP.NET MVC was an attempt to win Rails programmers back without feeling too enterprisey. So you had these very convenient, very frameworky solutions that maybe looked like you were leaning on the type system but really it was all just reflection. That became the standard in every language, and nobody needed to remember “parse don’t validate” because heavy frameworks did the work. And why not? Very few error or result types in fancy typed languages are actually suited for showing multiple (internationalised) validation errors on a web page. The bitter lesson of programming languages is that whatever clever, fast, safe, low-level features a language has, someone will come along and create a more productive framework in a much worse language. Note, this framework - perhaps the very last one - is now ‘AI’.
- jiehong 8mo agoAnd then clojure enters: let’s keep few data structures but with tons of method. So things stay as maps or arrays all the way through.
- renox 8mo agoI worked (a long time ago) on a C project where every int was wrapped in a struct. And a friend told me about a C++ project where every index is a uint8, uint16, and they have to manage many different type of objects leading to lots of bugs.. So it isn't really linked to the language.
- macintux 8mo agoA frequent visitor to HN. Tip: if you click on the "past" link under the title (but not the "past" link at the top of the page), you'll trigger a search for previous posts. https://hn.algolia.com/?query=Parse%2C%20Don%27t%20Validate&type=story&dateRange=all&sort=byDate&storyText=false&prefix&page=0 https://hn.algolia.com/?query=Parse%2C%20Don%27t%20Validate&... However, it's more effective to throw quotes into the mix, reduces false positives. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=%22Parse%2C%20Don%27t%20Validate%22&sort=byDate&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- pcwelder 8mo agoEach repost is worth it. This, along with John Ousterhout's talk [1] on deep interfaces was transformational for me. And this is coming from a guy who codes in python, so lots of transferable learnings. [1] https://www.youtube.com/watch?v=bmSAYlu0NcY https://www.youtube.com/watch?v=bmSAYlu0NcY
- curiousgal 8mo agoSemi tangent but I am curious. for those with more experience in python, do you just pass around generic Pandas Dataframes or do you parse each row into an object and write logic that manipulates those instead?
- lmeyerov 8mo agoPass as immutable values, and try to enforce schema (eg, arrow) to keep typed & predictable. This is generally easy by ensuring initial data loads get validated, and then basic testing of subsequent operations goes far. If python had dependent types, that's how i'd think about them, and keeping them typed would be even easier, eg, nulls sneaking in unexpectedly and breaking numeric columns When using something like dask, which forces stronger adherence to typings, this can get more painful
- adammarples 8mo agoSpeaking personally, I try not to write code that passes around dataframes at all. I only really want to interact with them when I have to in order to read/write parquet.
- whalesalad 8mo agoThe circumstances where you would use one or the other are vastly different. A dataframe is an optimized datastructure for dealing with columnar data, filtering, sorting, aggregating, etc. So if that is what you are dealing with, use a dataframe. The goal is more about cleaning and massaging data at the perimeter (coming in, and going out) versus what specific tool (a collection of objects vs a dataframe) is used.
- tomtom1337 8mo agoDefinitely do not parse each row into eg pydantic models. You lose the entire performance benefit of pandas / polars by doing this. If you need it, use a dataframe validation library to ensure that values are within certain ranges. There are not yet good, fast implementations of proper types in Python dataframes (or databases for that matter) that I am aware of.
- yakshaving_jgt 8mo agoI did a lightning talk on this topic last year, with a concrete example in Yesod. https://www.youtube.com/watch?v=MkPtfPwu3DM https://www.youtube.com/watch?v=MkPtfPwu3DM
- zdw 8mo agoThis is a great article, but people often trip over the title and draw unusual conclusions. The point of the article is about locality of validation logic in a system. Parsing in this context can be thought as consolidating the logic that makes all structure and validity determination about incoming data into one place in the program. This lets you then rely on the fact that you have valid data in a known structure in all other parts of the program, which don't have to be crufted up with validation logic when used. Related, it's worth looking at tools that further improve structure/validity locality like protovalidate for protobuf, or Schematron for XML, which allow you to outsource the entire validity checking to library code for existing serialization formats.
- jmholla 8mo agoWhen I came to this idea on my own, I called it "translation at the edge." But for me it was more that just centralizing data validation, it also was about giving you access to all the tools your programming language has for manipulating data. My main example was working with a co-worker whose application used a number of timestamps. They were passing them around as strings and parsing and doing math with them at the point of usage. But, by parsing the inputs into the language's timestamp representation, their internal interfaces were much cleaner and their purpose was much more obvious since that math could be exposed at the invocation and not the function logic, and thus necessarily, through complex function names.
- solomonb 8mo agoI disagree. I think the key insight is to carry the proof with you in the structure of the type you 'parse' into.
- zdw 8mo agoCould you clarify what you mean by "carry the proof"?
- solomonb 8mo agoFrom the article: validateNonEmpty :: [a] -> IO () validateNonEmpty (_:_) = pure () validateNonEmpty [] = throwIO $ userError "list cannot be empty" parseNonEmpty :: [a] -> IO (NonEmpty a) parseNonEmpty (x:xs) = pure (x:|xs) parseNonEmpty [] = throwIO $ userError "list cannot be empty" Both consolidate all the invariants about your data; in this example there is only one invariant but I think you can get the point. The key difference between the "validate" and "parse" versions is that the structure of `NonEmpty` carries the proof that the list is not empty. Unlike the ordinary linked list, by definition you cannot have a nil value in a `NonEmpty` and you can know this statically anywhere further down the call stack.
- danieltanfh95 8mo agoHot take: Static typing is often touted as the end all be all, and all you need to do is "parse, don't validate" at the edge of your program and everything is fine and dandy. In practice, I find that staunch static typing proponents are often middle or junior engineeers that want to work with an idealised version of programming in their heads. In reality what you are looking for is "openness" and "consistency", because no amount of static typing will save you from poorly defined or optimised-too-early types that encode business logic constraints into programmatic types. This is also why in practice alot of customer input ends up being passed as "strings" or have a raw copy + parsed copy, because business logic will move faster than whatever code you can write and fix, and exposing it as just "types" breaks the process for future programmers to extend your program.
- solomonb 8mo agoThis is such a tired take. The burden of using static types is incredibly minimal and makes it drastically simpler to redesign your program around changing business requirements while maintaining confidence in program behavior.
- danieltanfh95 8mo agoPeople keep saying this and yet in the decades of my career the industry bounces between being fully dynamic and fully typed according to the affordability of senior engineers. What you are saying are covered by tests, not types.
- yakshaving_jgt 8mo ago> What you are saying are covered by tests, not types. You know Haskell programmers write tests, right?
- jghn 8mo ago> I find that staunch static typing proponents are often middle or junior engineeers I wouldn't go this far as it depends on when the individual is at that phase of their career. The software world bounces between hype cycles for rigorous static typing and full on dynamic typing. Both options are painful. I think what's more often the case is that engineers start off by experiencing one of these poles and then after getting burned by it they run to the other pole and become zealous. But at some point most engineers will come to realize that both options have their flaws and find their way to some middle ground between the two, and start to tune out the hype cycles.
- kayo_20211030 8mo agoA great piece. Unfortunately, it's somewhat of a religious argument about the one true way. I've worked on both sides of the fence, and each field is equally green in its own way. I've use OCaml, with static typing, and Clojure, with maybe-opt-in schema checking. They both work fine for real purposes. The big problem arrives when you mix metaphors. With typing, you're either in, or you're out - or should be. You ought not to fall between stools. Each point of view works fine, approached in the right way, but don't pretend one thing is the other.
- r4victor 8mo agoIt seems modern statically-typed and even dynamically-typed languages all adopted this idea, except Go, where they decided zero values represent valid states always (or mostly). A sincere question to Go programmers – what's your take on "Parse, Don't Validate"?
- taylorallred 8mo agoNot speaking for all Go programmers, but I think there is a lot of merit in the idea of "making zero a meaningful value". Zero Is Initialization (ZII) is a whole philosophy that uses this idea. Also, "nil-punning" in Clojure is worth looking at. Basically, if you make "zero" a valid state for all types (the number 0, an empty array, a null pointer) then you can avoid wrapping values in Option types and design your code for the case where a block of memory is initialized to zero or zeroed out.
- masklinn 8mo agoOnly if you ignore the billion cases where it doesn't work, such that half the standard library explodes if you try to use it with zero values because they make no sense[0], special mention to reflect.Value's > Panic: call of reflect.Value.IsZero on zero Value And the "cool" stuff like database/sql's plethora of Null* for every single type it can support. So you're not really avoiding "wrapping values in Option types", you're instead copy/pasting ad-hoc ones all over, and have to deal with zero values in places where they have no reason to be, forced upon you by the language. And then of course it looks even worse because... not having universal default values doesn't preclude having opt-in default values. So when that's useful and sensible your type gets a default value, and when it's not it doesn't, and that avoids having to add a check in every single method so your code doesn't go off the rail when it encounters a nonsensical zero value. [0] or even when you might think it does, like a nil Logger or Handler
- r4victor 8mo agoThat's exactly the problem. Thanks for describing! What I find is people using linters to ensure all struct fields are initialized explicitly (e.g. https://github.com/GaijinEntertainment/go-exhaustruct https://github.com/GaijinEntertainment/go-exhaustruct), which is uhh...
- whalesalad 8mo agoThe author's point here is great, but the post does (imho) a poor job illustrating it. The tl;dr on this is: stop sprinkling guards and if statements all over your codebase. Convert (parse) the data into truthful objects/structs/containers at the perimieter. The goal is to do that work at the boundaries of your system, so that inside of your system you can stop worrying about it and trust the value objects you have. I think my hangup here is on the use of terms parse vs validate. They are not the right terms to describe this.
- tialaramex 8mo agoI understand where you're coming from, but these terms seem fine to me: This is exactly what, for example, Rust's str::parse method is for. The documentation gives the example: let four: u32 = "4".parse().unwrap(); You will so very often have text and want typed information, and parse is exactly how we do that transformation exactly once. Whereas validation is what it looks like when we try to make piecemeal checks later.
- lock1 8mo agoComing from a more "average imperative" background like C and Java, outside of compiler or serde context, I don't think "parse" is a frequently used term there. The idea of "checking values to see whether they fulfill our expectations or not" is often called "validating" there. So I believe the "Parse, Don't Validate" catchphrase means nothing, if not confusing, to most developers. "Does it mean this 'parse' operation doesn't 'validate' their input? How do you even perform 'validation' then?" is one of several questions that popped up in my head the first time I read the catchphrase prior to Haskell exposure. Something like "Utilize your type system" probably makes much more sense for them. Then just show the difference between `ValidatedType validate(RawType)` vs `void RawType::validate() throws ParseError`.
- tialaramex 8mo agoThe crucial design choice is that you can't get a Doodad by just saying oh, I'm sure this is a Doodad, I will validate later. You have to parse the thing you've got to get a Doodad if that's what you meant, and the parsing can fail because maybe it isn't one. let almost_pi: Rational = "22/7".parse().unwrap(); Here the example is my realistic::Rational. The actual Pi isn't a Rational number so we can't represent it, but 22 divided by 7 is a pretty good approximation considering. I agree that many languages don't provide a nice API for this, but what I don't see (and maybe you have examples) is languages which do provide a nice API but call it validate. To me that naming would make no sense, but if you've got examples I'll look at them.
- rorylaitila 8mo agoI make great use of value objects in my applications but there are things I needed to do to make it ergonomic/performant. A "small" application of mine has over 100 value objects implemented as classes. Large apps easily get into the 1000s of classes just for value objects. That is a lot of boilerplate. It's a lot of boxing/unboxing. It'd be a lot of extra typing than "stringly typed" programs. To make it viable, all value objects are code-generated from model schemas, and then customized as needed (only like 5% need customization beyond basic data types). I have auto-upcasting on setters so you can code stringly when wanted, but everything is validated (very useful for writing unit tests more quickly). I only parse into types at boundaries or on writes/sets, not on reads/gets (limit's the amount of boxing, particularly on reading large amounts of data). Heavy use of reflection, and auto-wiring/dependency injection. But with these conventions in place, I quite enjoy it. Easy to customize/narrow a type. One convention for all validation. External inputs are by default secure with nice error messages. Once place where all values validation happens (./values classes folder).
- metalliqaz 8mo agobonus points for the correct use of "cromulent"
- dang 8mo agoRelated. Others? Parse, Don't Validate (2019) - https://news.ycombinator.com/item?id=41031585 https://news.ycombinator.com/item?id=41031585 - July 2024 (102 comments) Parse, don't validate (2019) - https://news.ycombinator.com/item?id=35053118 https://news.ycombinator.com/item?id=35053118 - March 2023 (219 comments) Parse, Don't Validate (2019) - https://news.ycombinator.com/item?id=27639890 https://news.ycombinator.com/item?id=27639890 - June 2021 (270 comments) Parse, Don’t Validate - https://news.ycombinator.com/item?id=21476261 https://news.ycombinator.com/item?id=21476261 - Nov 2019 (230 comments) Parse, Don't Validate - https://news.ycombinator.com/item?id=21471753 https://news.ycombinator.com/item?id=21471753 - Nov 2019 (4 comments)
- macintux 8mo agoThere are other threads for articles inspired by it. Those with comments: Parsix - https://news.ycombinator.com/item?id=27166162 https://news.ycombinator.com/item?id=27166162 TypeScript - https://news.ycombinator.com/item?id=28425435 https://news.ycombinator.com/item?id=28425435 C - https://news.ycombinator.com/item?id=44507405 https://news.ycombinator.com/item?id=44507405 Without comments: Non-blank strings in Rust - https://news.ycombinator.com/item?id=34947030 https://news.ycombinator.com/item?id=34947030 Email type in Rust - https://news.ycombinator.com/item?id=34946791 https://news.ycombinator.com/item?id=34946791 Java - https://news.ycombinator.com/item?id=29250169 https://news.ycombinator.com/item?id=29250169
- LordDragonfang 8mo agoI'll be honest, as someone not familiar with Haskell, one of my main takeaways from this article is going down a rabbit hole of finding out how weird Haskell is. The casualness at which the author states things like "of course, it's obvious to us that `Int -> Void` is impossible" makes me feel like I'm being xkcd 2501'd.
- mrkeen 8mo agoIf you spend your life talking about bool having two values, and then need to act as if it has three or 256 values or whatever, that's where the weirdness lives. In C, true doesn't necessarily equal true. In Java (myBool != TRUE) does not imply that (myBool == FALSE). Maybe you could do with some weirdness! In Haskell: Bool has two members: True & False. (If it's True, it's True. If it's not True, it's False). Unit has one members: () Void has zero members. To be fair I'm not sure why Void was raised as an example in the article, and I've never used it. I didn't turn up any useful-looking implementations on hoogle[1] either. [1] https://hoogle.haskell.org/?hoogle=a+-%3E+Void&scope=set%3Astackage https://hoogle.haskell.org/?hoogle=a+-%3E+Void&scope=set%3As...
- tialaramex 8mo agoWhat were you expecting to find? A function which returns an empty type will always diverge - ie there is no return of control, because that return would have a value that we've said never exists. In a systems language like Rust there are functions like this for example std::process::exit is a function which... well, hopefully it's obvious why that doesn't return. You could imagine that likewise if one day the Linux kernel's reboot routine was Rust, that too would never return.
- mrkeen 8mo ago> What were you expecting to find? > functions like this for example std::process::exit
- tialaramex 8mo ago
- sevensor 8mo agoMaking illegal states unrepresentable sounds like a great idea, and it is, but I see it getting applied without nuance. “Has multiple errors” can be a valid type. Instead of bailing immediately, you can collect all of the errors so that they can be reported all together rather than forcing the user to fix one error at a time.
- mh2266 8mo agoIs this not `Result<Whatever, List<Error>>`? There's nothing enforcing that the error side needs to be the value-based equivalent of a single instance of an Exception class. The important part is not to expose a "String -> Whatever" function publicly.
- d0liver 8mo agoI think, more generally, "push effects to the edges" which includes validation effects like reporting errors or crashing the program. If you, hypothetically, kept all of your runtime data in a big blob, but validated its structure right when you created it, then you could pass around that blob as an opaque representation. You could then later deserialize that blob and use it and everything would still be fine -- you'd just be carrying around the validation as a precondition rather than explicitly creating another representation for it. You could even use phantom types to carry around some of the semantics of your preconditions. Point being: I think the rule is slightly more general, although this explanation is probably more intuitive.
- jmull 8mo agoSystems tend to change over time (and distributed nodes of a system don’t cut over all at once). So what was valid when you serialized it may not be valid when you deserialize it later.
- d0liver 8mo agoThis issue exists with the parsed case, too. If you're using a database to store data, then the lifecycle of that data is in question as soon as it's used outside of a transaction. We know that external systems provide certain guarantees, and we rely on them and reason about them, but we unfortunately cannot shove all of our reasoning into the type system. Indeed, under the hood, everything _is_ just a big blob that gets passed around and referenced, and the compiler is also just a system that enforces preconditions about that data.
- gaigalas 8mo ago> Now I have a single, snappy slogan that encapsulates what type-driven design means to me, and better yet, it’s only three words long IMHO this is distracting and sort of vain. It forces this "semantics" perspective into the reader, just so the author can have a snappy slogan. Also, not all languages have such freedom in type expressiveness. Some of them have but offer terrible trade-ofs. The truth is, if you try to be that expressive in a language that doesn't support it you'll end up with a horror story. The article fails to mention that, and that "snappy slogan" makes it look like it's an absolute claim that you must internalize, some sort of deep truth that applies everywhere. It isn't.
- waffletower 8mo agoI'm sorry, I don't like to title drop, but I am a Staff Data Engineer and I find that "type driven" development is an inappropriate world view for many programming contexts that I encounter. I use "world view" carefully as it makes a contractual assumption about reality -- "give me what I expect". Data processing does not always have the luxury of such imposition. In these contexts a dynamic and introspective world view is more appropriate, "What do we have here?" "What can we use?". In 2019 I would have felt crippled by use of Haskell in data processing contexts and have instead done much in Clojure in these intervening years, though now LLM assisted use of Haskell toward such tasks would be a fun spectator sport.
- yakshaving_jgt 8mo ago> I don't like to title drop, but I am a Staff Data Engineer I am a Chief Technology Officer[^1]. Your opinion here is common, and misguided. Here is why: https://lexi-lambda.github.io/blog/2020/01/19/no-dynamic-type-systems-are-not-inherently-more-open/ https://lexi-lambda.github.io/blog/2020/01/19/no-dynamic-typ... --- [^1]: Literally nobody cares.
- waffletower 8mo agoThat's an insular opinion piece that doesn't sway, especially in the age of AI agents, it has not aged well. Its shallow rejection of Rich Hickey's nuance, is also unconvincing. It is a polemical justification for a coding philosophy that is incomplete and dishonest about the benefits of alternatives. Thanks for reminding me that no one cares; important to reinforce that.
- yakshaving_jgt 8mo agoThat's quite the shallow dismissal, and the bit about AI agents is a particularly weird non sequitur — King's argument is about what type systems can and cannot express. AI agents don't change the relationship between static types and open-world data processing. It sounds like you're annoyed that Hickey's position was effectively challenged.
- 8mo ago
- Joel_Mckay 8mo agoAn unconstrained json/bson parser without recursive structure limits must be bounded somehow. In many cases, the ordering of marshaled data cannot be guaranteed across platforms. The best method is walk the symbolic tree with a cost function, and score the fitness of the data compared to expected structures. For example, mismatched or duplicate GUID/Account/permission/key fields reroute the message to the dead-letter queue for analysis, missing required fields trigger error messaging, and missing optional fields lower the qualitative score of the message content. Parsers can be extremely unpredictable, and loosely typed formats are dangerous at times. =3
- mmis1000 8mo agoThis article always end up relevant once in a while. Recently, I am trying to make llm to output specific format. It turns out no matter how you wrote propmt and perform validate. It will never be as effective as just limit the output with proper bnf (via llama cpp grammar file).
- 1-more 8mo agoA related talk is Richard Feldman's "Making Impossible States Impossible." Richard wrote a number of Elm packages and is the creator of the Roc language. https://www.youtube.com/watch?v=IcgmSRJHu_8 https://www.youtube.com/watch?v=IcgmSRJHu_8
- cbondurant 8mo agoA really mindset-altering read for me, I've carried this way of thinking ever since I'd first read it a few years ago.
- hackrmn 8mo agoThis article has done rounds on the ITernet before. Maybe because it resonates with people (who repost it time and again). Anyway, I very much agree with the idea. In my experience, "text" or "string" is not a type. Technically it is one, of course, but I seldom see good use of it for when a more apt type would do better -- in short, it's a last resort thing, and it fares badly there too. Ironically, the only good use for it is as input to a... parser. I see a lot of URLs being passed around as strings within a system perfectly capable of leveraging typing theory and offering user defined types, if not at least through OOP goodness a lot of people would furiously defend. The URL, in this case, would often have _already_ been parsed once, but effectively "unparsed" and keeps being sent around as text in need of parsing at every "junction" of the system that requires to meaningfully access it, except that parsing is approached like some ungodly litany best avoided and thus foregone or lazily implemented with a regex where a regex isn't nearly sufficient. Perhaps it's because we lack parsers, by and large, or in the very least parser generators that are readily available, understandable (to your average developer), and simple enough to use without requiring to understand formal language theory with Chomsky hierarchy, context sensitivity, grammar ambiguity and parse forests, to say the least. Same with [file] paths, HTTP header values, and other things that seem alluring to dismiss as only being text. It wouldn't be a problem, had I not seen time and again how the "text" breaks -- URLs with malformed query parameters because why not just do `+ '?' + entries.map(([ name, value ]) => name + "=" + value).join("&")`, how hard can it be? Paths that assume leading slash or lack there of etc. I believe the article was born precisely of the same class of frustrations. So I am now bringing the same mantra everywhere with me: "There is no such type as string". Parse at earliest opportunity, lazily if the language allows it (most languages do) -- breadth first so as to not pay upfront, just don't let the text slip through. I am talking from experience, really, your mileage may vary.
- exodys 8mo agoMaybe I am being contrarian, or maybe I don't understand; if I am reading input, I am always going to validate that input after parsing. Especially if it is from a user. I understand that they should be separate, but they should be very close together.
- Jtsummers 8mo ago> if I am reading input, I am always going to validate that input after parsing. In the "parse, don't validate" mindset, your parsing step is validation but it produces something that doesn't require further validation. To stick with the non-empty list example, your parse step would be something like: parse [h|t] = Just h :| t parse [] = Nothing So when you run this you can assume that the data is valid in the rest of the code (sorry, my Haskell is rusty so this is a sketch, not actual code): process data = do { Just valid <- parse data; ... further uses of valid that can assume parsing succeeded, if it didn't an error would already have occurred and you can handle it } That has performed validation, but by parsing it also produces a value that doesn't require any revalidation. Every function that takes the parsed data as an argument can ignore the possibility that the data is invalid. If all you do is validate (returning true/false): validate [h|t] = true validate [] = false Then you don't have that same guarantee. You don't know that, in future uses, that the data is actually valid. So your code becomes more complex and error-prone. process data = if validate data then use data else fail "Well shit" use [h|t] = do_something_with h t use [] = fail "This shouldn't have happened, we validated it right? Must have been called without data being validated first." The parse approach adds a guarantee to your code, that when you reach `use` (or whatever other functions) with parsed and validated data that you don't have to test that property again. The validate approach does not provide this guarantee, because you cannot guarantee that `use` is never called without first running the validation. There is no information in the program itself saying that `use` must be called after validation (and that validation must return true). Whereas a version of `use` expecting NonEmpty cannot be called without at least validating that particular property.
- exodys 8mo agoAh, I get it. So, it's just a tagging system. Once tagged, assume valid. DRY.
- throw567643u8 8mo agoWhat's lexi up to these days? Her last big contribution to Haskell was the delimited continuation primops, then she disappeared in a puff of smoke.
- sn9 8mo agoShe wrote about it in the most recent post on her blog: https://lexi-lambda.github.io/blog/2025/05/29/a-break-from-programming-languages/ https://lexi-lambda.github.io/blog/2025/05/29/a-break-from-p...
- tlavoie 8mo agoAlong with all the general discussion, I found the concept of defensive parsing striking a chord when reading this as well: "The Seven Turrets of Babel: A Taxonomy of LangSec Errors and How to Expunge Them", https://langsec.org/papers/langsec-cwes-secdev2016.pdf https://langsec.org/papers/langsec-cwes-secdev2016.pdf I'd love for these ideas to take hold at work, but I'm on the fringes in infosec, not a dev.
- benhoyt 8mo agoI'm not very familiar with functional programming and Haskell in particular. I think I understand the gist of this article, and "use data structures that make illegal states unrepresentable". However, is there a similar article but written with more common languages (C#, C++, Java, Go) in mind? Or is a big part of this concept only relevant for strong functional languages with sum types and pattern matching?
- masklinn 8mo agoIt is relevant to all languages with static type checkers from idris to python. But of course since it is about expressing properties via the type system the more expressive that is the easier and more applicable. Java has sum types, incidentally. And pattern matching.
- lock1 8mo ago> Or is a big part of this concept only relevant for strong functional languages with sum types and pattern matching? It need not strictly be a pure functional language for type-driven style to be usable. Type-driven style only requires the fact that some type cannot be assigned to another type, so it's kind of possible to do even in a language like C, as `int a = (struct Foo) {};` would get rejected by C compilers. However, I don't think it's doable in languages with structural type systems like Typescript or Go's interface without a massive ergonomic hit for minimal gain. Languages with a structural type system are deliberately designed to remove the intentionality of "type T cannot be assigned to type S" in exchange for developer ergonomics. > However, is there a similar article but written with more common languages (C#, C++, Java, Go) in mind? For C#, there's F#-focused article, which I believe some of it can be applied to C# as well: F# - Railway Oriented Programming - https://fsharpforfunandprofit.com/rop/ https://fsharpforfunandprofit.com/rop/ F# - Designing with Types - https://fsharpforfunandprofit.com/series/designing-with-types/ https://fsharpforfunandprofit.com/series/designing-with-type... For modern Java, there is some attempt at popularizing "Data-Oriented Programming" which just rebranded "Type-driven design". Surprisingly, with JDK 21+, type-driven style is somewhat viable there, as there is algebraic data type via `record` + `sealed` and exhaustive pattern match & destructuring. Inside Java Blog - Data-Oriented Programming - https://inside.java/2024/05/23/dop-v1-1-introduction/ https://inside.java/2024/05/23/dop-v1-1-introduction/ Infoq - Data-Oriented Programming - https://www.infoq.com/articles/data-oriented-programming-java/ https://www.infoq.com/articles/data-oriented-programming-jav... For Rust, due to the new mechanics introduced by its affine type system, there is much more flexibility in what you could express in Rust types compared to more common languages. Rust - Typestate Pattern - https://cliffle.com/blog/rust-typestate/ https://cliffle.com/blog/rust-typestate/ Rust - Newtype - https://rust-unofficial.github.io/patterns/patterns/behavioural/newtype.html https://rust-unofficial.github.io/patterns/patterns/behaviou...