10 ms·
Wrong way around: only allow http:// http:// and https:// https:// (and generally filtering out anything thats not letters, numbers, slash or dot is probably a
by hughperkins 10y ago
Wrong way around: only allow http:// http:// and https:// https:// (and generally filtering out anything thats not letters, numbers, slash or dot is probably a good idea. Remove any sequences of more than one slash or dot.
- mootothemax 10y ago>Wrong way around: only allow http:// http:// and https:// https:// For myself, it's a subtle change in developer thinking - "what should I allow" vs "what should I exclude" - that's paid off massively over the years.
- artursapek 10y agohttps://en.m.wikipedia.org/wiki/Robustness_principle https://en.m.wikipedia.org/wiki/Robustness_principle
- taneq 10y agoDidn't we decide in the aftermath of IE6 that this was a bad idea, and that we should be strict in both what we accept and what we emit?
- gsnedders 10y agoNo, we decided that what was important was interoperable implementations: it doesn't matter how you achieve that goal. What's needed is specs that define how to handle all input (it doesn't matter what the spec says: it can define how to handle every single last possible case as HTML5 does, or it can define a subset of inputs to trigger some fatal error handling as XML1.0 does) and sufficient test suites that implementers catch bugs in their code before it ships (and the web potentially starts relying on their quirks). The problem with IE6 was the fact that it wasn't interoperable (in many cases, every implementation was conforming according to the spec, and there were frequently differences in behaviour in valid input that the spec didn't fully define) and the fact that it had lots of proprietary extensions (and being strict and disallowing any extensions makes it hard to extend formats in general in a non-proprietary way; one option is strict versioning but then you end up with a million if statements all over the implementation to alter behaviour depending on the version of the content). Some of the worst issues with IE that took the longest for other browsers to match were things like table layout: IE quite closely matched NN4 having invested a lot in reverse-engineering that as the web depended on the NN4 behaviour in places; Gecko had rewritten all the table layout code from NN and didn't match its behaviour having been written according to the specs which scarcely define how to layout any table even today.
- mootothemax 10y ago>https://en.m.wikipedia.org/wiki/Robustness_principle https://en.m.wikipedia.org/wiki/Robustness_principle You need to be careful of where you place the emphasis on that, though: Be -liberal- in what you accept. vs Be liberal in what you -accept-.
- artursapek 10y agoRight - you accept URIs. That's fairly liberal. > Be conservative in what you do However, you only handle specific schemes and ignore the rest.
- mootothemax 10y agoYeah, apologies, I was being pretty petty. My hopefully-better-expressed point is that it's easy to interpret the robustness principle in different ways, some of which lead to better code, and some of which... don't.
- artursapek 10y agoNo worries. I think the philosophy is rooted in the fact that you can't control what other parties will send you; you can only control what you send in response. So that's the main thing to keep in mind. It's sort of like the good life advice you hear occasionally: you can't control other peoples' actions; only your own. Emotional maturity, etc.
- toomanybeersies 10y agoIsn't that missing the point of the robustness principle, which is more related to say, networking, and accepting things that aren't strictly to RFC spec, but when sending things, you match the spec to the letter?
- artursapek 10y agoI've found it useful in software design in general.
- myhf 10y agohttps://en.wikipedia.org/wiki/End-to-end_principle https://en.wikipedia.org/wiki/End-to-end_principle Most of the advice in this thread would accidentally disable ftp support.
- vetinari 10y agoAlso file:// urls pointing to server shares. (Anyway, only MSIE supports this from http(s) origin, and then people wonder, why MSIE is still being used).
- jerf 10y agoIt should be pointed out that while this was once accepted as gospel, it has been coming under a lot of fire lately. HTML, once arguably the flagship of this principle and its greatest success (I say "arguably" because you can also argue TCP), no longer works this way. HTML5 specifies how bad input should be handled, and if you accept that "how to process nominally bad input" as the "real" standard, HTML is now strict in what it accepts. It's just that what it is strictly accepting appears quite flexible. I'm not a big believer in it myself; "liberal in what you accept" and "comprehensible for security audits" are not quite directly opposed, but certainly work against each other fairly hard. There's a time and a place for Postel's principle, but I consider it more an exception for exceptional circumstances rather than the first thing you reach for.
- eyelidlessness 10y ago> HTML5 specifies how bad input should be handled, and if you accept that "how to process nominally bad input" as the "real" standard, HTML is now strict in what it accepts. HTML5 is a shining example of "be liberal in what you accept", and its improved documentation of how to handle bad input (note that bad input is still permitted!) greatly expands HTML's "be conservative in what you send". I think HTML5 is a perfect example of the Robustness Principle.
- jerf 10y agoThe "bad input" is, arguably, no longer bad input. The standard has been redefined to strictly specify what to do with that "bad" input, and if you don't handle it exactly as the standard specifies, it won't do what you "want" it to do. That's not "being liberal in what you accept". Being liberal in what you expect is what we had before HTML 5, where the standard specified the "happy case" and the browsers were all "liberal in what they expect", in different ways. I am not stretching any definitions here or making anything up, because "liberal in what you accept" behaviors in the real world demonstrably work this way; everybody is liberal in different ways. It can hardly be otherwise; it isn't "being liberal in what you accept" if you accept exactly what the standard permits, after all. When liberality is permitted, what happens in practice is that out-of-spec input is handled in whatever the most convenient way for the local handler is, in the absence of any other considerations (such as deliberately trying to be compatible with the quirky internal details of the competition). Browsers leaked a lot about their internal differences if you observed how they tended to handle out-of-spec input. Thus a standard like HTML5 that clearly specifies how to handle all cases now is fundamentally not "liberal in what it accepts" anymore. Instead, it is a rare, if not unique, example of a standard that has been rigidly specified after a couple of decades of seeing exactly how humans messed up the original standard. It is, nevertheless, now quite precise about what to do about the HTML you encounter. You aren't allowed to be "liberal", you're told exactly what to do.
- oldmanjay 10y agoNo, be strict, fail fast, and report the errors. Robustness is not achieved by muddling through on a misinterpretation, it is achieved by working toward correctness.
- artursapek 10y agoI think "fail fast" can work well in a closed and controlled system but when accepting input from many other parties it's not as practical or desirable.
- buro9 10y agoExactly. Whitelist only trusted schemes, do not wait to blacklist untrusted. I wrote the Go HTML sanitizer: https://github.com/microcosm-cc/bluemonday https://github.com/microcosm-cc/bluemonday and have a rule for user generated (untrusted) content that basically does whitelist just the things that one can trust: https://github.com/microcosm-cc/bluemonday/blob/master/helpers.go#L124-L141 https://github.com/microcosm-cc/bluemonday/blob/master/helpe... That states that URIs must be: 1. Parseable 2. Relative 3. Or one of: mailto http https 4. And that I will add rel="nofollow" to external links, and additionally I'll add "rel="noopener" if the link has a target="_blank" attribute Oh, and I do not trust Data URIs either.
- deleted 10y ago[deleted]
- stirner 10y agoMight want to add tel to the whitelist. It works in roughly the same way as mailto but interfaces with telephone apps instead of email clients.
- buro9 10y agoThis is the default user-generated policy, others are able to tweak and adjust using policy rules, i.e: p.AllowURLSchemes("tel") I chose conservative and safe defaults, not everyone wishes to whitelist telephone links.
- willvarfar 10y ago(Also be weary of imagetragick-type bugs too, where the URL starts innocuously and then contains some shellcode, because you pass the URL to something that'll paste it into system() call)
- throwanem 10y agoI certainly am weary of bug branding...
- willvarfar 10y agoI know, me too, but I didn't name it. There's a section on the main page imagetragick.com talking about branding and how they got no traction without one.
- nerdponx 10y agoI only started seeing this weary/wary misspelling in recent years. They don't sound alike, and they don't really look alike. Did cell phone spellcheckers give rise to this one?
- willvarfar 10y ago'weary' means tired of something, 'wary' means cautious of something.
- nerdponx 10y agoRight, and in the last year or so I've started to see people getting these terms confused.
- barefootcoder 10y agoBecause our language is a mashup of multiple other languages, and thus the rules are inconsistent. As for pronunciation of those two: weary is pronounced like "ear" wary is pronounced like "air"
- 10y ago
- pdkl95 10y agoCan we please stop trying to enumerate badness[1]? When parsing input it is possible to define the set of valid input, not all possible invalid inputs. Also, anybody accepting input from an untrusted source (such as anything from a network or the user) that isn't verifying the data with a formal recognizer is doing it wrong[2]. Instead of writing another weird machine, guarantee that the input is valid with a parser generator (or whatever) recognize the input and drop anything even slightly invalid. [1] http://www.ranum.com/security/computer_security/editorials/dumb/ http://www.ranum.com/security/computer_security/editorials/d... [2] https://media.ccc.de/v/28c3-4763-en-the_science_of_insecurity https://media.ccc.de/v/28c3-4763-en-the_science_of_insecurit...
- ptero 10y agoI agree with the approach. However, specific examples of different badnesses are useful for testing the final product.
- kazinator 10y agoI.e. move enumeration of badness from code into test-cases.
- ubernostrum 10y agoguarantee that the input is valid with a parser generator OK, that works really well... until you learn how much non-RFC-specified behavior is built in to web browsers. Simply building a parser to the RFC will leave you wide open to all sorts of nastiness! The is_safe_url() internal function in Django is a bit of a historical dive into things we've learned about how browsers interpret (or, arguably, misinterpret) various types of oddball URLs: https://github.com/django/django/blob/master/django/utils/http.py#L287 https://github.com/django/django/blob/master/django/utils/ht...
- oneplane 10y agoI do wonder; is there any browser that is actually full-RFC-specced? I checked a few (the mainstream desktop ones, but also links2 etc.), but so far they all seem to have glue to fix historical behavior.
- robert_tweed 10y agoJust as a concrete example of why this is the right approach, there is at least one enterprise CMS that installs a custom protocol handler that can be use to access any object stored in the CMS if you know or can guess/discover its URI. It's worth assuming there are others that you don't know about. The other advantage of the whitelist approach here is that you know exactly which protocols you think you support and can design tests for them. For instance to support https, you'll want to check you have decent error handling and do not silently accept potential MitM certificates.
- avian 10y ago> filtering out anything thats not letters, numbers, slash or dot is probably a good idea. This is highly non-trivial once you realize that the world speaks more than ASCII and things like http://www.xn--n3h.net http://www.xn--n3h.net exist.
- mootothemax 10y ago>This is highly non-trivial once you realize that the world speaks more than ASCII and things like http://www.xn--n3h.net http://www.xn--n3h.net exist. I was under the impression that requests to and from the server still used ASCII? That is, the server would see a host header as this: Host: www.xn--n3h.net And not as this: Host: www.[snowman icon].net Anything else is a question of URL-encoding, which if not used would raise interesting bugs with space characters, let alone anything more exotic like snowmen. Edit for completeness: in my server logs, the GET request for a /[snowman icon] URL is url encoded to GET /%E2%98%83 HTTP/1.1
- tokenizerrr 10y agoRight but how does the user submit it and what do you put in the href?
- kijin 10y agoIf a user copied the URL from the address bar, it will be correctly percent-encoded already. You can put the same percent-encoded URL in the href attribute of a hyperlink. A properly encoded URL will not contain any character that requires escaping in an HTML context. When a user clicks on that link, the browser will navigate to the percent-encoded URL but display the snowman icon in the address bar. If the user copies it, it will transparently turn back into the percent-encoded URL. All modern browsers do this.
- germanier 10y agoI just tried doing that with a few domain names containing an umlaut (äöü) and every single time that letter was copied into the clipboard (even though behind the scenes at the request level it would have been encoded). This is what I expect as a regular user. They don't want to deal with encoded, unreadable URLs.
- wodenokoto 10y agoI believe uri today accept all sorts of characters outside ascii.
- busterarm 10y agoBut if your code returns a URL, please don't do this. You should allow PRURLs (protocol-relative URLs). As in "//url". Especially you, Hubspot. I should be able to set a protocol-relative thank you page URL on your forms. If my user reaches your embedded form on my page as http, you should give them http. If they do it on https, it should give them https. Yes, I have an axe to grind.
- SnacksOnAPlane 10y agoYou seriously want to disallow gopher://? C'mon, man!
- deleted 10y ago[deleted]
- zwily 10y agoAlso make sure to fully resolve the DNS down to all possible IP addresses, and verify that they are all external to your network. And if you're on EC2, make sure nobody is hitting 169.254.169.254. Really, there are so many gotchas around fetching user-supplied URLs that it's scary.
- yuliyp 10y agoAnd be sure to check it again if there's a redirect. Don't let your URL handling library do this. Alternately send all your traffic through a proxy that can't talk into your network.
- l_zzie 10y agoImportantly, fetching DNS twice (once to check, another to download) is an incomplete solution, since DNS responses can change (cf "DNS rebinding").
- captn3m0 10y agoI remember this was exactly how a readability service (readability or instapaper or something similar, can't recall now) was attacked. The service allowed you to fetch internal urls and presented them formatted on your phone. A mixture of file:// and internal web urls allowed complete takeover.
- alxndr 10y agoAny chance you could dig up the details on that?
- 1propionyl 10y agoThis is the correct answer and best practice. Be conservative in what you accept and liberal in what you produce.
- zero_iq 10y agoAlso, fully decode the input string before doing this processing to make sure you really find all those sequences. A simple, seemingly obvious step that a surprising amount of software neglects to do.