5 ms·
But what if at some point somewhere down the line someone forgets to sanitize the output? Surely better procedure to sanitize at both ends. Nobody is perfect.
by bardan 7y ago
But what if at some point somewhere down the line someone forgets to sanitize the output? Surely better procedure to sanitize at both ends. Nobody is perfect.
- laumars 7y agoYou might need to escape strings differently depending on where you’re outputting eg HTML or JSON.
- cjfd 7y agoThis is completely the wrong attitude. Code should be correct, not fail-safe. All of the safety hatches that people tend to introduce make the code less predictable and eventual problems tend to arise far away from where they originated, making them difficult to diagnose. What is the result of this sanitization? Now we have some undefined, changing internal string format running around in our application, and possibly multiple undefined and changing internal string formats, where it is also unclear at what point a string is supposed to be in what format. If something is an arbitrary string it should be allowed to be an arbitrary string and the things handling that should escape it the appropriate way. The article is correct.
- Scarblac 7y agoCode should ideally be correct and fail-safe; deal gracefully with bad input, ensure correct output, at all levels. Ideally. That still doesn't mean you should care about output sanitization on data entry, as you don't know how it will be output yet.
- zAy0LfpBZLC8mAC 7y ago> Code should ideally be correct and fail-safe; deal gracefully with bad input, ensure correct output, at all levels. Ideally. NO! If you get bad input, you should fail, loudly. Anything else is a recipe for disaster. (Not as in "crash the whole system", of course, but as in "reject the request".)
- eythian 7y agoThis tends to leads to systems that don't work and die on real-world edge cases.
- zAy0LfpBZLC8mAC 7y agoNo, it doesn't. That is what "accept anything" programming leads to. The idea that "accepting all the inputs" somehow gives you an advantage is an illusion: If the semantics of some input are not well-defined, then the only thing you gain by accepting it anyway are hard to debug interoperability problems and vulnerabilities. When some input is not well-defined accoridng to the spec, then your interpretation is just a random guess, and the next developer will make a different random guess as to what that input means, and so an interoperability problem and potential vulnerability is born. If you reject the invalid input, you will notice the error and thus fix the source of the invalid input to produce input for which the semantics are actually well-defined.
- ashearer 7y agoThis sounds good in theory, but I'll give a counterexample. Requirement: Name input box. Implementation: We'll sanitize the input by rejecting any characters likely to be dangerous if mishandled, like single quotes, or anything else we don't immediately imagine to be useful. If a character turns out to be needed later, that's no problem. We'll just change the list. Security audit: Passes Later customer complaint: I can't sign up! — J. O'Brien Dev team: Sorry, too bad. We'd have to re-audit everything and possibly modify code to allow your last name, because there might be code somewhere that relies on the original sanitization for security. That was the point of sanitizing on input, after all. If you want to sign up, it would be easiest for us if you would just change your name.
- zAy0LfpBZLC8mAC 7y agoI think you misunderstood my point. I am not saying that you should reject valid (that is: semantically meaningful) input, but that if you are confronted with semantically meaningless input, you should reject it rather than garble it so that it gains some random meaning. So: Name input field, value "J. O'Brien": accept JSON parameter, value "{foo:bar}": reject The context was the idea that you should gracefully accept bad input. If your code considers "J. O'Brien" bad input for a name, then that's the problem, not that it doesn't accept bad input.
- XMPPwocky 7y agoDouble-escaping is silly & it\'s just plain incorrect.
- Scarblac 7y agoYou can't sanitize for output at input time, as the sanitization that needs to be applied is different for HTML, JS and JSON. You don't know that at input time.
- jve 7y agoWell, use libraries/frameworks that ENFORCE you to sanitize and makes that exceptional case to output raw content. Examples. PHP: Using mysql_escape_string is a no-no - you will forget to add it one day. Using parametrized queries you won't write unsafe SQL. .NET Core - Outputting to HTML by default only outputs those chars to HTML which are in predefined UTF range. All other chars will be converted to HTML entities. If you want to output raw, you must explicitly use @Html.Raw https://docs.microsoft.com/en-us/aspnet/core/mvc/views/razor?view=aspnetcore-3.1#expression-encoding https://docs.microsoft.com/en-us/aspnet/core/mvc/views/razor...
- JimDabell 7y agoYou're thinking about this as if data can be in one of two states – untrusted or sanitised. This is not the case. When you output arbitrary data, you need to encode it in a way that is suitable for that context. These contexts might be: - Generating a web page. - Including in a JSON response from an API. - Sending an email. - Storing in an SQL database. These all use different formats / protocols that use different syntax to encode data. How you correctly encode data for one of them is different to how you correctly encode data for another of them. There is no method of taking untrusted data and "sanitising" it so that it is correct for all of them. What works for one will break for the rest. If you want to handle arbitrary data correctly and safely, store it as-is and when the time comes to use it, encode it appropriately for the context you are using it in. Where possible, use tools and systems that get it right by default instead of requiring developers to remember to encode correctly, e.g. generate HTML with templating engines that encode data as HTML by default, and use parameterised queries with SQL.
- TeMPOraL 7y ago> generate HTML with templating engines that encode data as HTML by default Don't, unless you're sure the templating engine actually parses the HTML into a tree of nodes before interpolating and re-emitting it. Otherwise it's likely someone will interpolate something in an improper context, e.g. inside <script> or <style> block.
- thephyber 7y agoThis doesn't pass the sniff test. If someone can forget to sanitize the output, someone can also forget to sanitize the input. The most important things are to understand where the content is used, use the appropriate output encoding/escaping, have rigorous tests to ensure your expectation correctly escapes nasty strings, and that you keep the output escaping code up to date to protect against novel attacks and new browser/app features. I worked at a social media company with one of the largest text-based user-content-stores in the world at the time. Some of the features had input-side encoding and some had output-side encoding. I was there ~10 years after the bad practice of input-side encoding started and it very quickly became too cumbersome to know exactly which fields were encoded with what encoding (and I mean both character encoding and htmlentities / specialchars / specific character stripping / etc). We started getting ridiculous bugs like passwords could not contain '&' characters or logins would fail matching what we had in the DB. It's not about being perfect. That will never happen. It's about storing exactly what the user submitted (if it is accepted by the POST submission logic) and to correctly encode the output for the correct security context (HTML, XML, JSON, html entities, html attribute, script tag, styles/stylesheet, urls, uploaded filename / file contents, filesystem injection, command injection, etc). These all have different rules. You can unintentionally open yourself to a vulnerability in one if you only expect the output to be displayed in HTML.