8 ms·
In search of the perfect URL validation regex
- to3m 12y agoIf you're going to allow dotted IPs you should really allow 32-bit IPs too, e.g., http://0xadc229b7 http://0xadc229b7, http://2915183031 http://2915183031 and http://025560424667 http://025560424667. (The validity of this last one was news to me I must admit.)
- DanBlake 12y agoYep. I should mention that 99.9% of domains will fall into standard form ( handle.domain or ip.ip.ip.ip ) As such, You are definitely more likely to let a user enter a bad URL they did not intend because it validates then to let a uncommon domain actually be used. As such- a much simpler regex would likely 'make more people happy' than being 100% correct to tech spec.
- baddox 12y agoAlso IPv6 URLs, like http://[1080:0:0:0:8:800:200C:417A]/index.html http://[1080:0:0:0:8:800:200C:417A]/index.html http://www.ietf.org/rfc/rfc2732.txt http://www.ietf.org/rfc/rfc2732.txt
- gamegoblin 12y agoSince I just finished implementing a toy HTTP/1.1 server, I must throw in my newfound knowledge that 2732 has been updated with Zone IDs: http://tools.ietf.org/html/rfc6874 http://tools.ietf.org/html/rfc6874
- zAy0LfpBZLC8mAC 12y agoNone of those is a URI, so a URI validator most certainly should not accept them. Just because browsers tend to understand them as a matter of a historical accident does not mean those are valid URIs, just as tag soup that browsers also tend to understand isn't valid HTML either.
- tshadwell 12y agoI've put the test cases into a refiddle: http://refiddle.com/refiddles/53a736c175622d2770a70400 http://refiddle.com/refiddles/53a736c175622d2770a70400
- MatthewWilkes 12y agoWhy no IPv6 addresses in the test cases?
- VaucGiaps 12y agoWhy not put in some of the new TLDs as test cases... ;)
- siliconc0w 12y agoThis is a good lesson why you want to avoid writing your own regexes. Even something simple like an email address can be insane:http://ex-parrot.com/~pdw/Mail-RFC822-Address.html http://ex-parrot.com/~pdw/Mail-RFC822-Address.html
- baudehlo 12y agoThis always comes up in these discussions and is a terrible counter example. RFC 822 is the format for email messages (ie headers) and not a form you'll ever find email addresses in "in the wild" (eg on web forms).
- Dylan16807 12y agoWhy are http://www.foo.bar./ http://www.foo.bar./ and http://a.b--c.de/ http://a.b--c.de/ supposed to fail? The @stephenhay is just about perfect despite being the shortest. The subtleties of hyphen placement aren't very important, and this is a dumb place to filter out private IP addresses when a domain could always resolve to one. Checking if an IP is valid should be a later step.
- masklinn 12y agoThe first one I'm not sure, it looks like a valid FQDN (equivalent to www.foo.bar, where the trailing dot is implicit). The second one I guess it has to do with punycode and internaltional URI encoding?
- Dylan16807 12y agoYou'd only have to filter out xn--, though.
- saraid216 12y agoThe trailing dot makes it an invalid URL, which defines hostname as *[ domainlabel "." ] toplabel. "b--c" seems to be a valid domainlabel, though, so I'm not sure why it's on there.
- deleted 12y ago[deleted]
- thwarted 12y agoThe trailing dot makes it an invalid URL, which defines hostname as ∗[ domainlabel "." ] toplabel. RFC2396§3.2.2 defines hostname as: hostname = *( domainlabel "." ) toplabel [ "." ] RFC3986 which obsoletes RFC2396, seems to make fewer claims about the authority section, and specifically says in Appendix D.2 that the toplabel rule has been removed.
- LukeShu 12y ago> The trailing dot makes it an invalid URL, which defines hostname as * [ domainlabel "." ] toplabel. I disagree. But, this is a tricky one. The relevant specs are: Spec | | Validity | Definition URL (RFC1738) | obsolete | invalid | hostname = *[ domainlabel "." ] toplabel HTTP/1.0 (RFC1945) | current | invalid | host = <A legal Internet host domain name | | | or IP address (in dotted-decimal form), | | | as defined by Section 2.1 of RFC 1123> HTTP/1.1 (RFC2068) | obsolete | invalid | ; same as RFC1945 HTTP/1.1 (RFC2616) | obsolete | valid | hostname = *( domainlabel "." ) toplabel [ "." ] URI (RFC3986) | current | valid | host = IP-literal / IPv4address / reg-name | | | reg-name = *( unreserved / pct-encoded / sub-delims ) HTTP/1.1 (RFC7230) | current | valid | uri-host = <host, see [RFC3986], Section 3.2.2> The only way that URL is invalid is if we are in a strict HTTP/1.0 context. As a note about RFC1738 being obsolete: these days a URL is just a URI (1) whose scheme specifies it as a URL scheme, and (2) is valid according to the scheme specification. As the given URL is a valid URI, and is valid according to the current http URL scheme specification (RFC7230), that URL is valid.
- deleted 12y ago[deleted]
- eli 12y agoAt best this lets you conclude that a URL could be valid. Is that really useful? Is the goal here to catch typos? Because you'd still miss an awful lot of typos. If you really want your URL shortener to reject bad URLs, then you need to actually test fetching each URL (and even then...) As an aside, I'd instantly fail any library that validates against a list of known TLDs. That was a bad idea when people were doing it a decade ago. It's completely impractical now.
- masklinn 12y ago> At best this lets you conclude that a URL could be valid. Is that really useful? It's useful to find and linkify URLs in text (e.g. in your HN comments, how do you think HN makes http://foo.com http://foo.com into a link?)
- eli 12y agoThat's not the premise laid out on the linked page.
- mathias 12y agoMy exact use case was the following: the user clicks a bookmarklet that passes the current URL in the browser as a query string parameter to a URL shortener script. The validation is then performed before the URL is shortened. In that scenario, and with the given requirements, I can’t think of a case where the validation fails. There’s no need to worry about protocol-relative URLs, etc. (Keep in mind that this page is 4 years old — I very well may have missed something.) > If you really want your URL shortener to reject bad URLs, then you need to actually test fetching each URL (and even then...) I disagree. http://example.com/ http://example.com/ might experience downtime at some point in time, but that doesn’t mean it’s suddenly an invalid URL. > As an aside, I'd instantly fail any library that validates against a list of known TLDs. That was a bad idea when people were doing it a decade ago. It's completely impractical now. Agreed.
- eli 12y agoI still don't quite follow the purpose of the validation. Is it against malicious use? In normal use, I would think that pretty much any URL that's good enough for the browser sending it would be good enough for the link shortener.
- eridius 12y agoJohn Gruber (of daringfireball.com) came up with a regex for extracting URLs from text (Twitter-like) years ago, and has improved it since. The current version is found at https://gist.github.com/gruber/249502 https://gist.github.com/gruber/249502. I haven't tested it myself, but it's worth looking at. Original post: http://daringfireball.net/2009/11/liberal_regex_for_matching_urls http://daringfireball.net/2009/11/liberal_regex_for_matching... Updated version: http://daringfireball.net/2010/07/improved_regex_for_matching_urls http://daringfireball.net/2010/07/improved_regex_for_matchin... Most recent announcement, which contained the Gist URL: http://daringfireball.net/linked/2014/02/08/improved-improved-regex http://daringfireball.net/linked/2014/02/08/improved-improve...
- masklinn 12y agoIsn't it the @gruber v2 column from the page? It looks to have no false negative, but many false positives. The only one which does perfect on the tested set is Diego Perini's https://gist.github.com/dperini/729294 https://gist.github.com/dperini/729294
- eridius 12y agoHrm, you're right. I managed to miss seeing that while skimming the page.
- bdarnell 12y agoAnother important dimension when evaluating these regexes is performance. The Gruber v2 regex has exponential (?) behavior on certain pathological inputs (at least in the python re module). There are some examples of these pathological inputs at https://github.com/tornadoweb/tornado/blob/master/tornado/test/escape_test.py#L20-29 https://github.com/tornadoweb/tornado/blob/master/tornado/te...
- keeperofdakeys 12y agoFrom experience, the python re module does weird things sometimes. There is a better third-party regex module, https://pypi.python.org/pypi/regex https://pypi.python.org/pypi/regex.
- pipeep 12y agoDoes it use NFA? http://swtch.com/~rsc/regexp/regexp1.html http://swtch.com/~rsc/regexp/regexp1.html Because the issue with the URL regex mentioned is with backtracking.
- pipeep 12y agoIn node.js too. I found this out the hard way. I ended up modifying it so that it didn't work as well, but at least stopped DoSing my service: https://github.com/PiPeep/NotVeryCleverBot/blob/coffee-rewrite/src/transform/indexer.coffee#L58 https://github.com/PiPeep/NotVeryCleverBot/blob/coffee-rewri... Note the commented out lines in the here-regex.
- lucb1e 12y agoWhat's wrong with IP-address URLs? If they are invalid because it says so in some RFC, this is still not the ultimate regex. If you redirect a browser to http://192.168.1.1 http://192.168.1.1 it works perfectly fine. And why must the root period behind the domain be omitted from URLs? Not only does it work in a browser (and people end sentences with periods), the domain should actually end in a period all the time but it's usually omitted for ease of use. Only some DNS applications still require domains to end with root dots.
- mcpherrinm 12y agoThis is the context of a URL shortener, and one of the goals seems to prohibit bad IPs, not all IPs. Bad seems to be defined here as ones in the private IP spaces, and those in multicast space, etc.
- lucb1e 12y agoOkay, that makes sense!
- gamegoblin 12y agoAccording to the RFC, since domains are allowed to be entirely numeric, there is overlap between valid domains and valid IP addresses. The RFC says that if something could be a valid IP address, it is to be thought of as an IP address.
- timmm 12y agoWhat flavor of regex are we do making this in?
- mdavidn 12y agoUse a standard URI parser to break this problem into smaller parts. Let a modern URI library worry about arcane details like spaces, fragments, userinfo, IPv6 hosts, etc. uri = URI.parse(target).normalize uri.absolute? or raise 'URI not absolute' %w[ http https ftp ].include?(uri.scheme) or raise 'Unsupported URI scheme' # Etc
- droope 12y agoIt'd be interesting to look at URI.parse's source code.
- mdavidn 12y agoIt does match a regular expression internally. https://github.com/ruby/ruby/blob/trunk/lib/uri/rfc2396_parser.rb https://github.com/ruby/ruby/blob/trunk/lib/uri/rfc2396_pars...
- mathias 12y agoThat was not an option in this case, as the goal is to validate URLs entered as user input and blacklist certain URL constructs even though they’re technically valid.
- mdavidn 12y agoSo check if uri.host matches your blacklist.
- TazeTSchnitzel 12y agoWhy use a regex? It's much simpler to write a URL validator by hand, speaking as someone who wrote a URL parser,[1] and fixed a bug in PHP's.[2] Or, you know, use a robust existing validator or parser. Like PHP's, for instance. [1] https://github.com/TazeTSchnitzel/Faucet-HTTP-Extension https://github.com/TazeTSchnitzel/Faucet-HTTP-Extension - granted, this deliberately limits the space of URLs it can parse, but it's not difficult to cover all valid cases if you need to [2] https://github.com/php/php-src/commit/36b88d77f2a9d0ac74692a679f636ccb5d11589f https://github.com/php/php-src/commit/36b88d77f2a9d0ac74692a...
- 6cxs2hd6 12y agoExactly. Isn't this as bad an idea as trying to parse HTML with regular expressions? [1] [1]: http://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags/1732454#1732454 http://stackoverflow.com/questions/1732348/regex-match-open-...
- zAy0LfpBZLC8mAC 12y agoNo, it's not, URIs are regular, so using regular expressions is perfectly fine.
- 6cxs2hd6 12y agoGood point. You're right, it's not as bad an idea. (I still think it's not a great idea. Being regular isn't necessarily the same as being parseable with a maintainable regex.)
- dep_b 12y agoI agree, the correct one (+500 chars) looks like a maintenance nightmare to me. I tried to build a real e-mail address validator that would accept also the more exotic forms and there was no way in hell I would have done that in regex. I also met very few people that actually understood regex. It's a whole new skill and if I use it in my application I don't know if the next guy can pick it up.
- droope 12y agoI just validate with this regex '^http' :P
- Sir_Cmpwn 12y agoWhen you have a hammer, everything looks like a nail.
- lazyloop 12y agoYou do realize that RFC 3986 actually contains an official regular expression, right? http://tools.ietf.org/html/rfc3986#appendix-B http://tools.ietf.org/html/rfc3986#appendix-B
- TazeTSchnitzel 12y agoThat's for parsing a "well-formed" URI, not validating a URI as well-formed in the first place, however :)
- saraid216 12y agoThat doesn't work for validation.
- lazyloop 12y agoOf course it does, if that regular expression matches we have a valid URI. It's silly what gets downvoted on HN these days.
- zAy0LfpBZLC8mAC 12y agoNo, it doesn't. This regular expression matches strings that are not URIs, and that should be quite obvious if you compare the grammar in that same RFC to the regex.
- lazyloop 12y agoYes, it does. And it is quite obvious if you compare the grammar to the regular expression.
- zAy0LfpBZLC8mAC 12y agoPlease show how to produce the following string (which is matched by that regex) from the grammar (without the quotes): "%x" (edit: actually, feel free to do it with the quotes included if you like, that would still be matched by that regex)
- cobalt 12y agoWhat's wrong with /([\w-]+:\/\/[^\s]+)/gi It's not fancy but it will essentially match any url
- zAy0LfpBZLC8mAC 12y ago/./ will also match any URL. The point is to reject non-URLs.
- mathias 12y agoYes, to reject non-URLs and also some URLs that are technically valid but that I want to explicitly disallow anyway.
- JetSpiegel 12y agoIt has to match this valid URL: http://موقع.وزارة-الاتصالات.مصر http://موقع.وزارة-الاتصالات.مصر
- zAy0LfpBZLC8mAC 12y agoWTF? When will people finally learn to read the spec and implement things based on the spec and test things based on the spec instead of just making up themselves what a URL is or what HTML is or what an email address is or what a MIME body is or ... There are supposed URIs in that list that aren't actually URIs, there are supposed non-URIs in that list that are actually URIs, and most of the candidate regexes obviously must have come from some creative minds and not from people who should be writing software. If you just make shit up instead of referring to what the spec says, you urgently should find yourself a new profession, this kind of crap has been hurting us long enough. (Also, I do not just mean the numeric RFC1918 IPv4 URIs, which obviously are valid URIs but have been rejected intentionally nonetheless - even though that's idiotic as well, of course, given that (a) nothing prevents anyone from putting those addresses in the DNS and (b) those are actually perfectly fine URIs that people use, and I don't see why people should not want to shorten some class of the URIs that they use.) By the way, the grammar in the RFC is machine readable, and it's regular. So you can just write a script that transforms that grammar into a regex that is guaranteed to reflect exactly what the spec says.
- akerl_ 12y agoThis entire rant begins with the premise that the spec matches the real world implementation. Given that one of your examples is "what an email address is", I submit that expecting reality and the spec to match is a beautiful dream, from which a developer should awaken before trying to implement such a scheme.
- zAy0LfpBZLC8mAC 12y agoExcept there is no "the real world implementation". There only are lots of implementations that are incompatible with the spec as well as amongst each other. Inventing yet another variant of your own that also isn't going to be compatible with anything is not going to help anyone. Deviating from the formal spec because everyone practically agrees how to do things, albeit differently than in the formal spec, is something quite different from making shit up, and actually tends to be even harder than building things to spec, as there tends to be no easy reference to look things up in, but instead you might have to look into the guts of existing implementations and talk to people who have built them to figure out what to do - and you would normally start with an implementation according to spec anyhow, and only add special cases for non-normative conventions lateron. Also, what exactly is the problem with email addresses? There is a very unambiguous grammar of those in the RFC, and there are lots of implementations of exactly what the spec specifies. Just because some web kiddies have made up some shit about email addresses and use that for validation, doesn't mean that postfix, qmail, or exim are written by morons.
- Buge 12y agoInterestingly it seems http://✪df.ws http://✪df.ws isn't actually valid, even though it exists. ✪ isn't a letter[1], so it isn't allowed in international domain names. I was looking at the latest RFC from 2010 [2] so maybe it was allowed before that. The owner talks about all the compatibility trouble he had after he registered it [3]. The registrar that he used for it, Dynadot, won't let me register any name with that character, nor will Namecheap. [1] http://www.fileformat.info/info/unicode/char/272a/index.htm http://www.fileformat.info/info/unicode/char/272a/index.htm [2] http://tools.ietf.org/html/rfc5892 http://tools.ietf.org/html/rfc5892 [3] http://daringfireball.net/2010/09/starstruck http://daringfireball.net/2010/09/starstruck
- X-Istence 12y agoIt's the same way with http://💩.la/ http://💩.la/ it gets turned into the following using IDNA: xn--ls8h.la. That is a valid domain name and should be treated as such.
- Buge 12y agoI guess you could argue the definition of "valid". According to the RFC it's DISALLOWED: Those that should clearly not be included in IDNs. Code points with this property value are not permitted in IDNs.
- CMCDragonkai 12y agoWhat does the red vs green boxes mean?
- CMCDragonkai 12y agoOh I get it now. Got confused between 1s and 0s.
- mnot 12y agoThere is no perfect URL validation regex, because there are so many things you can do with URLs, and so many contexts to use them with. So, it might be perfect for the OP, but completely inappropriate for you. That said, there is a regex in RFC3986, but that's for parsing a URI, not validating it. I converted 3986's ABNF to regex here: https://gist.github.com/mnot/138549 https://gist.github.com/mnot/138549 However, some of the test cases in the original post (the list of URLs there aren't available separately any more :( ) are IRIs, not URIs, so they fail; they need to be converted to URIs first. In the sense of the WHATWG's specs, what he's looking for are URLs, so this could be useful: http://url.spec.whatwg.org http://url.spec.whatwg.org However, I don't know of a regex that implements that, and there isn't any ABNF to convert from there.