4 ms·
Agree and Disagree. Sanitization has it's place, but from a user perspective it's better to just outright reject (through validation) inputs that aren't valid.
by InitialBP 5y ago
Agree and Disagree. Sanitization has it's place, but from a user perspective it's better to just outright reject (through validation) inputs that aren't valid.
There are often unexpected ways that data gets into the system (IT manually adding data, internal support tool to help customers add data, etc.) You need to ensure that you're properly sanitizing your input at every single input faucet and your sanitization has to predict how, when, and where it will be used by sanitizing for dangerous characters in filenames, shell, spreadsheet formula vulns, and XSS attacks.
Instead, (Or In addition to) just make the assumption that data in the database is dangerous, and ensure that you properly escape for your use case when using that data.
Using a username to create a new file? Escape for filenames based on which OS/language your using.
Using birthdates in an excel file? Escape for excel formulas.
Using bio on an HTML page? HTML Escape.
Using username as part of a URL path? URL Escape.
And finally circle back to the fact that sanitization where you change user input without their knowledge (like the "O'brien" -> "Obrien" example in the article) creates for a frustrating user experience.
- hamilyon2 5y agoI agree, when your app does exporting, use escaping and be happy. Nobody ever challenged that. But that is not enough. You should do defence in depth. What I am talking about, you can't realistically escape for every use, because 1) once it is stored, it is usually outside of your control. You simply do not know where your data will end up, due to e.g. new integrations that will be developed in future. 2) you can even not know the proper escaping rules for document types you are producing due to software obscurity. Nobody I can think of escapes any csv files for excel-2001 vulnerabilities. This is just one exaple of software where those files can actually end up opened. What is more economical/rational to change, your input validation or every csv/excel exporter/converter ever in existence?
- InitialBP 5y agoYour argument against escaping is the same one against input sanitization. Literally you cannot know where all your data will go when storing it so how do you even begin to sanitize that data? No special characters? Alphanumerics only? You're going to have a bunch of modified data that user's are unhappy with to start. Furthermore you have the same problem with simply not knowing all of the ways that data will get into your system, so you won't be able to reliably sanitize all of your inputs. (e.g. manual entry, internal tools, random scripts, etc.) I understand defense in depth and by all means doing some sanitization up front can help increase your defensive posture (Often at the cost of user experience), but the proper and effective way to protect any of 100000s of use cases for your data is to ensure that each use case is escaping data relevant to that use case. Direct responses: 1. If you store simply store the data, it is not your responsibility to ensure that the data is "sanitized", but simply to ensure that the data follows the expected format. Rather *the onus of security is on the person actually Using the data.* 2. I have tested applications and written up findings where customer's use internal data to generate excel files w/ PII in them, and one user could steal another the PII of employees from another customer due to CSV injection. Yes these attacks are weird and niche, and even if this one attack doesn't matter, it's just an example of various weird attacks. If you are writing software where you take user data and put it into strange filetypes or systems, you should be researching those systems and writing code that doesn't break or have unexpected issues. You can find the necessary technical documents to figure out how to escape things. (https://owasp.org/www-community/attacks/CSV_Injection https://owasp.org/www-community/attacks/CSV_Injection) Finally, input validation is not input sanitization. Validation is making sure it conforms to your expected data-type (e.g. Username cannot contain special characters.) Sanitization is when the application modifies data that the user has submitted - e.g. stripping off special characters from names (O'Brien => OBrien). You keep saying "Do defense in depth" but your argument is why sanitization is better. Do both, but ALWAYS escape if you have to choose.
- hamilyon2 5y agoThank you for good answer. English is not my main language, it is fully my fault if i didn't make my position clear. I fully agree with everything you said. Always escaping output is the right thing. Sanitization is not what you want, like, ever. > Literally you cannot know where all your data will go when storing it so how do you even begin to sanitize that data? You construct a minimum regular language your inputs should confirm to. In programming parlance this is called regex. This requires knowledge of your domain and, yes, a bit of extra work. So, the comment field on payment becomes an alphanumeric with certain unicode characters. A pain for a user who wants to send a code examples, sure. But should code examples be in payment comment field? And so on. Identifier field? Allow only latin letters and numbers. Custom css user wants to use? Restrict to valid css. Name of some goods to sell? First character is alphabetical, rest alphanumeric and [-#"'*|:,] but never two consecutive special characters, even with spaces between. Pdf? Construct a reasonable pdf regex and match for it. You should publish your regexp and guarantee your database contents should adhere to it. I think this is practical and reasonable thing to do. It would fail against a determined sophisticated attacker, but at least it might give it hard time, and make attack a bit more detectable and a bit less effective.