14 ms·
Don’t try to sanitize input, escape output (2020)
- wnoise 5y agoCompare with https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate/ https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
- dnautics 5y agoI think the domain models are different. "parse-don't-validate" is great when your users are internal and trusted (e.g. a library that does codegen - the operators of the parser are already in the codebase). When your users are potentially hostile, you should at some level have a separate validate and eject strategy.
- mjw1007 5y agoI think "parse, don't validate" is an improvement on what the author of this article recommends: « So in cases where you do need to “echo” raw user input, carefully filter input based on a restrictive whitelist, and store the result in the database. When you come to output it, output it as stored without escaping. » I think the "parse, don't validate" approach comes out as follows: - take the list of things you would have included on your whitelist - add nodes for them to your internal representation for parsed markdown - extend your markdown parser to convert html-like input into those nodes - implement output for those nodes in a similar way to normal markdown This way, given the "escape output" they recommend, it's harder for any variant of the input that you hadn't considered to have harmful effects.
- yakshaving_jgt 5y agoNo, the parse, don’t validate idea is completely unrelated. It’s about leveraging a type system like in Haskell. It’s about parsing a value into a type with a narrower domain which in turn minimises the amount of control flow needed to implement a sufficiently correct program.
- mjw1007 5y agoI agree with "It’s about parsing a value into a type with a narrower domain", but I don't see how you get to the first sentence. In their example of a markdown renderer, the internal-representation node is the type with the narrower domain.
- yakshaving_jgt 5y agoThe idea of parsing over validation is just as applicable with untrusted input as with trusted input. The idea is more about system design rather than the prevention of security vulnerabilities.
- dnautics 5y agocome again? I didn't say you can't validate while parsing for untrusted input, I said for untrusted input you will probably STILL need additional separate validation methods. Key emphasis on negating don't absolute imperative in the original aphorism.
- yakshaving_jgt 5y agoI didn't say that that's what you said. I tried to communicate that whether or not the input is trusted is besides the point. Unless I'm deeply confused, the idea in parse, don't validate is about doing something like this parseFoo :: Text -> Maybe Foo parseFoo t = if textIsAFoo t then Just (Foo t) else Nothing f :: Foo -> IO () f = _ Rather than something like this (which seems to be more common) f :: Text -> IO () f t = when (textIsAFoo t) g where g = _
- dnautics 5y agoI thought the idea of parse, don't validate is that your parser should contain validation logic. So instead of text -> generic json parser -> validate json -> ... you would do text -> custom json parser that stops if encounters "incorrect content" -> ...
- yakshaving_jgt 5y agoYou can think of the parser as containing validation logic — it can parse the value into a more constrained type if it conforms to validation rules, or it can fail. The point is that once your value is in a more principled type, the rest of the system is free from having to make assumptions (and guard against potential failures) about the breadth of that value's domain. As the article mentions, this is really only relevant to languages with proper type systems like Haskell.
- marcosdumay 5y agoTwo different advises for two different things. That one is about data validation, making sure it is coherent and fits your data quality rules. This one is about data encoding, making sure it fits a different system's rules.
- kevincox 5y agoThese are both good advise. I have seen really funny bugs where Java accepted non-ascii numbers in an IP address but the C++ control plane very much did not. If the re-serialized version was sent to the backend this wouldn't have been an issue. But the domains are different. Data validation is ensuring that the information is something that your system accepts. Data encoding is used when you are serializing information. You should very likely validate on input, but not "sanitize" or encode. You do your encoding on output.
- Ingaz 5y agoArticle starts from "don't sanitize input" but in the middle : > If you’re not using Markdown but want to let your users enter HTML directly, you only have the second option – you must filter using a whitelist. So "don't filter but filter"
- joering2 5y agoEvery online form where user can interact and send data back to a server is always a nightmare in terms of security. I do utilize mod_secure, but with my next project, I have an idea of doing "base64" on everything in client's browser via javascript then sending it to server and checking on backend if content is a valid base64. Is that a good concept?
- lesquivemeau 5y agoWouldn't prevent XSS afaik
- justinsaccount 5y agoThat could work if you are just going to store things as base64. It accomplishes nothing if you are going to decode the base64 on the backend and then use the original value as-is. If anything it's worse than nothing, because now mod_secure will just see the base64 content and might fail to detect certain attacks.
- afavour 5y agoUnfortunately that wouldn't help with a whole lot. The danger with input is that it could be used to e.g. escape a SQL query and delete your database. Which is why we now have parameterised queries and such to help alleviate those worries. If you think about it the process you're describing already happens: the browser sends the user's input as (usually) UTF8 string data, then the server decodes it. Changing that process to base64 wouldn't change much.
- deleted 5y ago[deleted]
- adrr 5y agoWouldn’t base64ing your inputs bypass mod_security?
- joering2 5y agook thank you everyone for your responses (+1s) - I was research on this idea and couldn't find anything online - now I know why!
- iou 5y agoDo both pls.
- hombre_fatal 5y agoIf you're doing both, I'd ask you what you think you're accomplishing by sanitizing input, especially when you're already escaping output. All you're doing is corrupting the data with a ritual that seems like it's securing something, and it tends to make you think that your data is now ready to be rendered anywhere without issue.
- jerf 5y agoI can't emphasize this enough. This isn't a matter of taste, like, maybe you sanitize, maybe you escape on the way out, it's all good, it all works, it's just a matter of opinion. Sanitizing the input is wrong. Actively, objectively, unrecoverably wrong. Once you've destroyed your data you can't get it back. Huge amounts of effort have been wasted by people trying to fix and recover data that was destroyed by systems "helpfully" "sanitizing" data. God help you if you have a sequence of these systems in a row each doing their own "sanitization" before you get the data. Do not "sanitize" your inputs. Do not tell other developers to sanitize their inputs. Do not sagely spout off on HN about the importance of sanitizing your inputs. It is wrong. The only "sanitization" that should be done is that when encoding to the output there are sometimes things that should simply be removed. For instance, a good HTML escaping function probably ought to entirely drop nulls, not even encoding them as or anything, just drop them. Some of the other ASCII characters are straight-up illegal in HTML as well, even encoded. But all that sort of "sanitization" should be in the escaping step. If you want to reject null characters at input time, that's part of validation, not sanitization.
- asplake 5y agoValidate inputs, escape outputs
- Buttons840 5y ago
- gkoberger 5y agoThis solution doesn't match the problem. Even the SQL injection example shows him sanitizing the input, which is at odds with the title of the post. Log4J is a more recent example of it being too late/useless to escape the output.
- brodouevencode 5y agoYeah little Bobby Droptables is still a thing.
- marcosdumay 5y agoWhat? SQL injection is avoided at the point of usage. Trying to sanitize your input against it is an extremely bad practice. The same is true about HMTL injection (whether you call it XSS or something else). Log4j is an example of not interpreting text that the developer was never aware that was code. It's kinda of the extreme opposite of escaping your text on usage.
- gkoberger 5y agoThe article says DON'T sanitize when putting it into the database. I think contextual escaping counts as "sanitizing input", so the solution of "don't try to sanitize input" is undermined.
- shawnz 5y agoI interpreted the message as not sanitizing inputs at the point they are received, a la PHP magic quotes. Instead, escape at the output (the output to the database engine).
- gkoberger 5y agoNo where in the article do they use "output" to mean from the database engine; they use it to mean "outputting HTML".
- 5y ago
- 1970-01-01 5y ago¿Por qué no los dos?
- pwdisswordfish9 5y agoBecause if someone wants to register with the name O'Malley, you should not refuse them, or worse, mangle their name.
- Avamander 5y agoThen that sanitation is incorrect. I don't think this discussion would has any merit if we're speaking about incorrect implementations.
- pjerem 5y agoOk, and imagine we are on an Internet forum, talking about what is correct sanitization, and I want to make the following example : <script>alert(42)</script> Will HN remove the <> characters and make my comment incomprehensible or will it escape it on output, preserving all the meaning ? (Well, I’ll know after hitting Reply) Edit : good boy, HN
- nybble41 5y agoYou can't sanitize "correctly" if you don't know where the data will be used. This is exactly why the article advocates for escaping output (e.g. immediately before inserting a string into a SQL query) rather than sanitizing input (e.g. by deleting single-quotes or other potentially problematic characters from strings as soon as they're received).
- 1970-01-01 5y agoIf someone wants to register with the username O'Malley and password O`Malley, would you let them?
- simonw 5y agoThat's not sanitization, that's validation. You implement a validation rule that says that the password and the username can't be the same thing, then use that to redisplay the registration form with an error message.
- ffhhj 5y agosanitize (client side) => confirm with user => trim+escape (server side) => insert
- dagss 5y agoWhy "escape"? Just insert. Using SQL parameters.
- InitialBP 5y agoInsert using SQL parameters is "escaping". The parameterization ensures that the data being passed gets interpreted by the DB as the expected data type by ensuring special characters aren't interpreted as "special" in that context.
- dagss 5y agoI think that is a strange use of the word "escape". You "escape" from something, in this context the query string. If parameters are not passed inside the query string then how can you say they are escaped? At least for the database I am familiar with (mssql), the query string is one parameter in the binary protocol, and then there are the other parameters that are not the query string which are used as arguments. Your usage here is a bit like saying that a standalone PNG file is "escaped" from the HTML document it is referenced from...since it is marked as not being HTML...
- InitialBP 5y agoAlternatively... validate (client side) => insert using sql parameterization (escaping) => escape per context when outputting Sanitizing is the idea that you are cleaning dangerous things from the original input (different than validating which is disallowing user's to input characters that don't conform to what your program expects). One BIG issue here is that validation is generally clear to the user ("That is an invalid email address") whereas sanitization normally doesn't consult or inform the user that there were changes and may result in unexpected things happening from a user perspective. From article: name is "John O'Brien" now displays as "John OBrien" (this is a trivial example but still an issue) The name thing is a great example of things you might not expect your users to do but are still totally valid use cases. Sanitization can be Extremely frustrating from a user perspective.
- parhamn 5y agoIt's cool to see how these posts are becoming less and less important in the wake of today's frameworks/tools protecting devs by default. From ORMs escaping SQL, to FE frameworks escaping html/js, to browsers starting to default to same-site=lax. It feels like we've slowly pulled ourselves out of OWASP hell. Pretty nice to see! Obviously it's still important (see log4j) to know it all especially when its not so clear cut, but still good progress.
- erosenbe0 5y agoI think we really failed in earlier eras to get it right due to the momentum of the frameworks. I would liken to some of the crap building materials that were allowed in the past as new, cheap alternatives but subsequently showed failure or hazards after short service-lifes. Contractors were tasked with implementing these materials to stay within budget and everyone suffered the effects later.
- swlkr 5y agoA strong content security policy also helps with xss
- ipaddr 5y agoInstead of sanitizing input you create unsafe datastore which might be used in other applications later. Do it as soon as possible.
- frontiersummit 5y agoI think it cuts both ways, as anyone who has needed to mine an existing data set for a new purpose can attest. Having the data sanitized can may your parsing job infinitely easier, while it can simultaneously destroy data which would have been extremely helpful to the new project.
- ipaddr 5y agoIf it doesn't fit into a data standard you are enforcing, it shouldn't exist in the database. There is nothing wrong with capturing the original text in a field or separate table.
- ncc-erik 5y agoI think what makes this hard for folks is tracking what the expected form of data is at each step of its lifecycle, especially considering people working with new and unfamiliar codebases or splitting focus on multiple projects. There are some frameworks that try using types to solve the problem. Alternatively, the developers could throw in a comment that looks something like: // client == submits raw data ==> web_server == inserts raw data (param. sql stmt) ==> db_server ==> returns query with raw data ==> our_function == returns html-escaped data ==> client
- nostrademons 5y agoI think a better way to think of this may be in terms of canonicalization. Inside your application, you should decide on a single canonical way to represent data, one which fits the type of processing and expected use of the application. For example, you might decide that all strings should be UTF8, and should be interpreted (and stored) as whatever the user initially wrote. You might decide that any structured data should be parsed and then stored as protobufs in a BigTable. Or you might decide that an RDBMS is your native datastore and use whatever the native string encoding is for it, as well as parse & normalize data into tables upon input. Then, whenever you take input, your job is to validate and encode it. If you get a Windows-1252 string, you should re-encode it to utf8 for further storage. If it has data that are invalid UTF-8 codepoints, you should either strip, replace with a replacement character, or notify the user with a validation failure. Same with structured data that fails your normalization rules - you should usually notify the user. And when you send output, you should escape based on the intended output device. If you're putting it in an HTML page, HTML-escape it. If it's a URL, url-encode it. If it's a database query, SQL escape it. If it's a CSV, quote it. Thinking in these terms keeps the internal logic of your application simple (there are no format conversions except at system boundaries), and it also gives you a lot of flexibility to preserve the user's intent and add new output formats later.
- platz 5y agoso you would prevent stored XSS attacks by escaping on the output step instead of the canonicalization step
- simonw 5y agoRight - the way to avoid XSS is to escape on output. Most good template languages these days implement auto-escaping of variables that are interpolated into HTML. You still have to be careful embedding content into non-HTML contexts. One classic example there is outputting a blob of JSON inside a <script> tag - you need to make sure that you handle the case where a string could contain "</script><script>evil_code_here()</script>".
- 5y ago
- dang 5y agoDiscussed at the time: Don’t try to sanitize input – escape output - https://news.ycombinator.com/item?id=22431022 https://news.ycombinator.com/item?id=22431022 - Feb 2020 (280 comments)
- chriswarbo 5y agoThe fundamental problem is attempting to conflate a bunch of semantically-distinct things, just because they might happen to (sometimes) be represented in memory by similar byte sequences. Such 'byte coincidences' lead to lazy, non-sensical operations, like "append this user-provided name to that SQL statement"; implemented by munging together a bunch of bytes, without thought for how they'll be interpreted. A much better solution is to ignore whether things might just-so-happen to be represented in a similar way in memory; and instead keep things distinct if they have different semantic meanings (like "name", "SQL statement", "HTML source", "shell command", "form input", etc.). That way, if we try to do non-sensical things like appending user input to HTML, we'll get an informative error message that there is no such operation. This isn't hard; but it requires more careful thought about APIs. Unfortunately many languages (and now frameworks) have APIs littered with "String"; ignoring any distinctions between values, and hence allowing anything to be plugged into anything else (AKA injection vulnerabilities)
- deleted 5y ago[deleted]
- AtNightWeCode 5y agoNo, garbage in, garbage out. Sure, things like log or SQL injections should not only be solved by sanitizing. You solve it by separating data and code. A lot of times you really want to store data in a structured canonical way. Usernames for instance. It is bad if you with Unicode trickery can create multiple usernames that looks the same. Product descriptions, it is bad if your ML needs to handle HTML and so on.
- kevincox 5y agoThis is wrong. If I leave a comment `'; DROP TABLE users; --` You should display it back in the app as exactly that. If you put it into an HTML attribute you escape the `'` and if you stick it in SQL you use parametrized statements. There is nothing "wrong" with that initial input. What is wrong is pasting it into an SQL string, HTML element, HTML attribute, URL parameter or anywhere else without properly encoding it. This is the main reason you can't "sanitize" input. You need to know what the output format is to properly encode it. There are different requirements if you are pasting it into a sed replacement command vs HTML attribute vs HTML element body. You can strip everything except a-zA-Z and cross your fingers but even that isn't necessarily sufficient for all output formats.
- ehutch79 5y agousing parameterized statements is sanitizing inputs into the database.
- kevincox 5y agoThe database is "outside" of your application server. You communicate with the database using statements and when you get the value back from the database it is unchanged. The encoding was just for transfer, no data has actually been changed.
- AtNightWeCode 5y agoMaybe a better way to put is that you should be smart about why, when, and where to sanitize your data. A comment on a forum should not remove “‘; DO BAD THINGS;”. Why would it? It is just text in probably some UTF8 encoding. No viable web framework will write it out in a raw format if you do not explicitly ask for it. In SQL you use parameters. But as I wrote in my original comment. There are several scenarios and if you work with a web, probably the most cases, where you really want to make sure that what you have stored is a clean structured canonical data representation. Not only for your security but also for third party consumers and analyzing. I understand that everybody who sells NOSQL solutions disagree.
- hamilyon2 5y agoSanitizing inputs is not what you realistically want. You should prohibit certain types of input. Whitelisting strings is that what I would call it. You should escape outputs, of course (not that anyone in 2022 thinks otherwise). Why escaping outputs alone won't work is because user inputs will be stored in some database and you can't realistically predict how, when, where it will be used. Years in the future. User name could be used as a filename once, opening up possibility of shell-based exploit. It could trigger a little-known spreadsheet formula vulnerability when exported for analysis. Novel, interesting xss attacks are common and produced every day. That could be even not your code, but the code your client or partner organisation run. You just never know. One common defence is user names (and other freeform fields) should not be allowed to be arbitrary bytes. That is defence in depth, an established practice.
- wongarsu 5y agoThat works well for things you can limit to alphanumeric, which is pretty much only usernames. For everything else there will be an exploit in some context without proper escaping. You can decrease the attack surface, but you have to weigh that against the false sense of security it might give developers.
- InitialBP 5y agoAgree and Disagree. Sanitization has it's place, but from a user perspective it's better to just outright reject (through validation) inputs that aren't valid. There are often unexpected ways that data gets into the system (IT manually adding data, internal support tool to help customers add data, etc.) You need to ensure that you're properly sanitizing your input at every single input faucet and your sanitization has to predict how, when, and where it will be used by sanitizing for dangerous characters in filenames, shell, spreadsheet formula vulns, and XSS attacks. Instead, (Or In addition to) just make the assumption that data in the database is dangerous, and ensure that you properly escape for your use case when using that data. Using a username to create a new file? Escape for filenames based on which OS/language your using. Using birthdates in an excel file? Escape for excel formulas. Using bio on an HTML page? HTML Escape. Using username as part of a URL path? URL Escape. And finally circle back to the fact that sanitization where you change user input without their knowledge (like the "O'brien" -> "Obrien" example in the article) creates for a frustrating user experience.
- whoopdedo 5y agoSounds like a restatement of Postel's robustness principle[1]. Did it go out of style to "be conservative in what you send, be liberal in what you accept" and we need to relearn it again? Well, perhaps it did. History has shown the dangers of not handling malformed input well. Postel's principle has received scrutiny[2] for reinforcing those mistakes by creating a mistaken belief in robustness. More recent recommendations have been to be stricter in handling of inputs[3]. But I think there is some confusion between robustness and defensiveness. "Be liberal in what you accept" may be confused with "don't sanitize your inputs" when not sanitizing is the less liberal action. Robustness means the program should not fail if it receives input it didn't expect. A program that crashes, hangs, executes unintended shell code, mangles the data, changes the thermostat, or other undefined behavior is not being robust. To prevent that from happening then data must be sanitized at input so that it can be processed without those side-effects. The examples of programs failing robustness have been because they were insufficiently defensive. The bigger issue is that robustness doesn't scale easily. You may know how your bit of code will deal with malformed data, but what about every other library you use? Or other systems you communicate with? It becomes a backstage problem, where once someone has gained access to a restricted area it's assumed they are authorized to be there. The further down the tech stack you go the less likely the code will be defensive. That puts a burden on the public-facing sanity checks to anticipate how relaxed they can be about the input. If you change the definition of output to include internal-outputs, then Postel's principle gets new life. That is, try not to program the entire system and ecosystem at once, but treat each software component as an island. Be liberal not only with the data you receive from the end-user, but also with return values from functions. Be conservative and escape not only your generated HTML, but also the SQL statements you dispatch to the backend. This is what input sanitizing is actually about, it's keeping the promise to the other parts of your program that your code isn't going to give them bad data. That's also what the linked article is saying, because the HTML being generated is itself one component in a chain of programs that includes the end-user's browser. [1] https://en.wikipedia.org/wiki/Robustness_principle https://en.wikipedia.org/wiki/Robustness_principle [2] https://programmingisterrible.com/post/42215715657/postels-principle-is-a-bad-idea https://programmingisterrible.com/post/42215715657/postels-p... [3] https://datatracker.ietf.org/doc/html/draft-iab-protocol-maintenance https://datatracker.ietf.org/doc/html/draft-iab-protocol-mai...
- blibble 5y agoguess I'll just put that 2gb "first name" directly into my database then
- pornel 5y agoThat's validation, not sanitization.
- Sebb767 5y ago> The parallel for SQL injection might be if you’re building a data charting tool that allows users to enter arbitrary SQL queries. You might want to allow them to enter SELECT queries but not data-modification queries. In these cases you’re best off using a proper SQL parser [...] to ensure it’s a well-formed SELECT query – but doing this correctly is not trivial, so be sure to get security review. If you are ever in this situation, you should actually use a dedicated read-only user that can only access the relevant data. If you need to hide columns, use views. Trying to parse SQL can easily go very wrong, especially when someone (ab-)uses the edge cases of your DB.
- gumby 5y agoSince you don't know where your output will end up how could you possibly know the syntax to escape it? And how can the consumer of an arbitrary string trust that every input will have been properly escaped?
- pornel 5y agoYou can't escape it ahead of time, for the same reason you can't reliably block or remove "dangerous" inputs ahead of time — you can't reliably know all the places and contexts they will be used in. So you escape at the point of use, as late as possible, when you know exactly what escaping you need. It's also easy to forget to escape. This is why it's best to have tools and practices that automate it, e.g. HTML templating engine that escapes everything by default, e-mail composing library that automatically converts text to whatever MIME magic is required, etc.
- scotty79 5y agoI'm really surprised by the discussion here. It's so obviously true and I realized this when correct php function to escape string for sql was names mysql_real_escape_string
- billpg 5y agoShameless plug: NEVER Sanitize Your Inputs (by me, 2013) https://billpg.com/never-sanitize-your-inputs/ https://billpg.com/never-sanitize-your-inputs/
- taneq 5y agoI think escaping output is making the same mistake as sanitizing input. What we should really be saying is "stop using string interpolation/concatenation to process generic user data". By default, text should only ever be treated as a blob. Yes, there are circumstances where it needs to be treated otherwise but they should be seen as a giant flashing 'danger' sign indicating the need to go back to sanitizing etc.
- Sohcahtoa82 5y agoEvery time this topic comes up, the comments are full of people talking past each other because they're operating under different definitions of "sanitize", "input", and "escape". And now in this case, we add "output" to the confusion. Is the SQL query you send to your DB input or output?