3 ms·
I'll bite - why is 2 right but 4 wrong? Is it because it's not reversible? Is it in case your sanitization rules change?
by beepy 5y ago
I'll bite - why is 2 right but 4 wrong? Is it because it's not reversible? Is it in case your sanitization rules change?
- lucideer 5y agoFirstly, to distinguish sanitization -vs- validation: sanitization retains information lossily, validation throws. Validation may increase security, but the primary motivator is stability (avoiding unexpected state) and UX (relevant error messaging). Input sanitization is intended purely as a security measure but is at best insufficient and usually reduces system security. The first problem is that input sanitization is sanitizing without context: you don't know what threats you're fighting because threat vectors of this variety target the output format. Will your input variable be included into a HTML template, an Atom feed, saved to some file that's parsed later, sent in a JSON API, templated into a style or script block, used in an SQL query. The possibilities aren't known at input (or if they are they will change as your application grows) so protecting against all possibilities isn't viable. Often people html-escape or blindly html-strip their inputs, regardless of whether they'll ever be used in a html template. Given the above you might think input sanitization is at worst useless, inefficient, but not harmful to security, but that brings us to the follow-up problems: lossiness & unknown state. If you're doing input sanitization, you're not doing input validation (at least not properly). Input sanitization is about accepting lossy values (threats removed) when you receive unexpected input. That means you're accepting unexpected input, which leads to all the problems input validation is designed to protect against. Finally, there is of course double-escaping. As mentioned above, people often html-escape as part of input sanitization. This essentially disables your ability to do reliable secure output escaping because that leads to double escaped values in output (or horrible double-escape-reversal hacks in output templating code). Generally you want to be securely output/escaping everything, which means you want a system where you can rely on always receiving clean unescaped values into your output template. I have encountered too many systems using a good, secure html templating library with output escaping baked in by default, where devs had to disable output escaping explicitly because some inputs were already pre-escaped. In theory you can keep track of which vars need raw output and which vars don't, but that's an unmanageable mess in practice. And its impossible to automate enforcement.
- sirsinsalot 5y agoOr in other terms: Store Raw, output safe means you can improve output safety over time. Storing safe locks you in to that mutation.
- lucideer 5y agoThat's a really nice succinct way of putting it. Will steal.
- sirsinsalot 5y agoJust don't steal my weird capitalisation hah