7 ms·
Do both pls.
by iou 5y ago
Do both pls.
- hombre_fatal 5y agoIf you're doing both, I'd ask you what you think you're accomplishing by sanitizing input, especially when you're already escaping output. All you're doing is corrupting the data with a ritual that seems like it's securing something, and it tends to make you think that your data is now ready to be rendered anywhere without issue.
- jerf 5y agoI can't emphasize this enough. This isn't a matter of taste, like, maybe you sanitize, maybe you escape on the way out, it's all good, it all works, it's just a matter of opinion. Sanitizing the input is wrong. Actively, objectively, unrecoverably wrong. Once you've destroyed your data you can't get it back. Huge amounts of effort have been wasted by people trying to fix and recover data that was destroyed by systems "helpfully" "sanitizing" data. God help you if you have a sequence of these systems in a row each doing their own "sanitization" before you get the data. Do not "sanitize" your inputs. Do not tell other developers to sanitize their inputs. Do not sagely spout off on HN about the importance of sanitizing your inputs. It is wrong. The only "sanitization" that should be done is that when encoding to the output there are sometimes things that should simply be removed. For instance, a good HTML escaping function probably ought to entirely drop nulls, not even encoding them as or anything, just drop them. Some of the other ASCII characters are straight-up illegal in HTML as well, even encoded. But all that sort of "sanitization" should be in the escaping step. If you want to reject null characters at input time, that's part of validation, not sanitization.
- asplake 5y agoValidate inputs, escape outputs
- Buttons840 5y agoYes, but remember in a lot of cases nearly anything is valid input.
- talideon 5y ago_Some_ sanitisation is fine. For instance, stripping leading and trailing space in some fields, case normalisation, automatic insertion of spaces in credit card numbers, that kind of thing. That is to say, you should sanitise as an affordance to the user. Given the choice between presenting an error to the user and automatic sanitisation, the latter is preferable. It's something that should be done carefully, but it's still good. Thoughtless sanitisation is a whole different kettle.
- jerf 5y agoI agree that cleanup is acceptable, and there's certainly some wiggle room in what people call cleanup vs. sanitization and such. But when people chant "sanitize your inputs" and expect it to be treated as sage wisdom, it's in a security context, and it is wrong in that context. Sanitization is not a valid security tool. Mind you, you might be forced into it if your back is against the wall and you're working on other code that is broken and you can't fix that other code's broken failure to escape or whatever. But it's still wrong, just a wrong thing you were forced to do. A richer point of view is more "don't destroy data you don't 100% mean to destroy". Whitespace in the wrong place or stray nulls can meet that bar. Removing characters for "security" reasons doesn't. Destroying data to prevent security issues downstream is not a good idea.
- nybble41 5y agoTo me that sounds more like canonicalization than sanitation. Depending on your requirements it might be fine to convert the input to a canonical form before processing. If you do this, be certain to do it before validation so that you don't accidentally "canonicalize" validated input into something which wouldn't pass the validation checks. A key aspect of canonicalization compared to sanitation is that the result should be something that the user would consider equivalent to their original input. The most common offender in my experience is the abuse of case normalization, especially for data like email addresses which are not defined as case-insensitive (at least for the mailbox name) even if many servers treat them that way. If you don't preserve the original case (and other parts such as "+" labels whose meaning is defined by the mail server) the address may not work at all, or may result in sending messages to the wrong user. Names, as an intimate part of the user's identity, are another area where case normalization can sometimes prove annoying or even offensive. If some legacy system requires names to be entered as all-caps US-ASCII characters, fine, but at least don't turn "O'Conner" or "MacDouglas" into "O'conner" or "Macdouglas" in some misguided attempt to ensure that just the first letter is capitalized. (And in some situations the first letter shouldn't be capitalized, e.g. the "dos Santos" in "Giovani dos Santos Ramírez"[0]—which is a single surname, not two names.) [0] https://en.wikipedia.org/wiki/Giovani_dos_Santos https://en.wikipedia.org/wiki/Giovani_dos_Santos
- serious_habit 5y agoIf I'm reviewing code and someone is implementing escaping that's an immediate, massive, red flag. It's SO HARD to get right and there are many MANY libraries for doing it correctly. The scary thing is how many bugs still make it into these libraries. Strongly prefer using an established library and see designs such as https://web.dev/trusted-types https://web.dev/trusted-types.
- AnonHP 5y ago> Sanitizing the input is wrong. Actively, objectively, unrecoverably wrong. I agree on the “unrecoverably” (sic) part, but strongly disagree on words like “objectively”. It can be bad only if the input sanitization is poorly done. If that’s poorly done, then it’s also likely that the output sanitization may be poorly done. One cannot then say that output sanitization is objectively bad because someone doesn’t know or care enough to do it properly. This is a complex topic that deserves more attention, not hand waving away with claims that cannot stand on their own.
- unclebucknasty 5y agoThe downside of not sanitizing inputs is that the data may well end up in contexts or new apps altogether wherein escaping output is not reliably done or known to be necessary. In an ideal world it shouldn't happen. But, it does, especially given turnover within organizations and the long lifespans of many systems and their datasets, lack of documentation, etc. So, this calls into question the viability of a strict policy for every scenario that encourages blithely storing <script>doEvil()</script> in the database. Seems some level of sanitization has a place as an extra layer of defense for some use cases.
- serious_habit 5y agoEven better- never sanitize your data. You should only use templating systems which safely handle user data. Don't use innerHTML assignments, don't concatenate user data into SQL queries. Use existing, validated libraries for generating HTML and SQL.
- JxLS-cpgbe0 5y ago
- pydry 5y ago>If you're doing both, I'd ask you what you think you're accomplishing by sanitizing input, especially when you're already escaping output. https://en.m.wikipedia.org/wiki/Defence_in_depth_(non-military) https://en.m.wikipedia.org/wiki/Defence_in_depth_(non-milita...
- hombre_fatal 5y agoI'd argue that sanitization makes things worse from that standpoint. What exactly was transformed in some given data and for what context? What needs to be done to reverse the sanitization process if you want to see the verbatim data, if that's even possible? Now that you want to escape the output, how can you reverse the sanitization transform so that you aren't double-escaping? What were the assumptions being made when this data was sanitized and what was that transform? In other words, it's simpler to hold the verbatim data and then ask "ok, how does it need to be escaped for this context?" than having to ask that same question with arbitrarily mangled data while worrying if the data was sufficiently escaped for this context at input-time some point in the past. Even beginners get almost all mileage from parameterized SQL queries + using an HTML templating library that escapes by default which is almost all of them these days. I think knee-jerk sanitization is a relic of the days where that wasn't common, namely <?php echo $username ?>, which wasn't necessarily the worst advice when you otherwise had to remember to echo htmlEscape($username) every single time. Fortunately, things have improved since those days.
- pydry 5y agoI've used a bunch of sanitizers and never had any issues with any of them. I'm sure there are exceptions but IME they tend to mangle the kind of text which the user really has no legitimate need to enter most of the time. Far from being a relic the recent log4j vulnerability highlighted just how much value there is in this kind of defense in depth. Obviously knee jerk decisions in tech are usually bad news.
- AnonHP 5y agoThe data store may be one, but the teams and apps working on the inputs and the outputs may be disparate and different. Relying on other teams all the time to do things correctly may not be a wise approach.
- deleted 5y ago[deleted]