5 ms·
A student and I have been using coverage-guided grammar-aware differential fuzzing to discover bugs in URL parsers for a while now. There is extreme variation i
by bkallus 3y ago
A student and I have been using coverage-guided grammar-aware differential fuzzing to discover bugs in URL parsers for a while now. There is extreme variation in this space; it's trivial to turn up meaningful bugs in widely-used URL parsers.
".://" is a particularly egregious example. (and, by the same principle, "evil.com://good.com")
- Python 3.6's urllib.parse sees the "." as the URL's scheme, and an empty authority.
- Python 3.11's urllib.parse sees the entire ".://" as the URL's path.
- urllib3.util.parse_urlsees the "." as the URL's hostname, the ":" as the separator for an empty port number, and the "//" as the path. (this is one of the most downloaded packages on PyPI)
- Boost::URL rejects the URL outright.
If you're going by RFC 3986, then only Boost::URL is exhibiting the correct behavior.
If you're going by the WHATWG URL standard, then I don't know which one of these behaviors (if any) is correct.
If you're interested in collaborating on this project, please send me email. My address is in the footer at https://kallus.org https://kallus.org
- gmac 3y agoAs another example of an egregious difference that might bite you, compare this JS in recent Node vs web browsers: new URL('postgres://user:pass@host/db') Node parses this as if it were a web address, into protocol, username, password, host, pathname, etc. Browsers just parse out the protocol and call everything else a pathname.
- btilly 3y agoBrowsers do that for a good reason. They used to do the protocol, username, password, host, pathname, etc. But scammers used it to have a user name that looked like a well-known URL, while actually directing the user to a domain under the scammer's control. Not honoring the spec was therefore a security feature.
- cma 3y agoWhy not just delete the url bar contents after attempting that url or something? Or take to a warning page.
- btilly 3y agoAny such decision requires no longer honoring https://www.rfc-editor.org/rfc/rfc1738 https://www.rfc-editor.org/rfc/rfc1738.
- cxr 3y agoPlease be more specific.
- btilly 3y agoSection 3.1 specifies /<user>:<password>@<host>:<port>/<url-path> Which means if you're presented with that and don't send it to that host, you've violated the RFC. But this was demonstrably resulting in confused users being sent to scammer's domains.
- cxr 3y agoThat ("if you're presented with that and don't send it to that host, you've violated the RFC") has nothing to do with the comment you responded to, which described a UI choice compatible with RFC 3986, which says "Applications should not render as clear text any data after the first colon (":") character found within a userinfo subcomponent" (and which also goes on to say "Applications may choose to ignore or reject such data when it is received as part of a reference").
- silverwind 3y ago`URL` in node and browser is supposed to be compatible, this looks like a bug (browser ignoring protocols they don't know).
- cxr 3y agoURL parsing semantics are defined by and dependent upon the scheme. (That's spec.) By definition, if you don't recognize the scheme, you cannot guarantee a correctly parsed URL. From RFC 1738: The Internet Assigned Numbers Authority (IANA) will maintain a registry of URL schemes. Any submission of a new URL scheme must include a definition of an algorithm for accessing of resources within that scheme and the syntax for representing such a scheme. The behavior described (extracting the scheme and treating the rest as an opaque string) is pretty much the only thing you can do when the scheme is unrecognized. (The other options being to throw an exception or return null.) Based on the description, it sounds like neither are breaking spec—it's just that Node supports "postgres". That is, unless it's true that Node's URL implementation is supposed to match what browsers do, in which case Node is breaking spec—its own.
- bkallus 3y agoRFC 3986 provides a generic URI grammar that is not scheme-specific, though other standards that define URL schemes may choose to subset the subset that grammar as they see fit. If a URL parser does not recognize a scheme, I would expect it to parse the URL using the generic parsing procedure.
- cxr 3y agoWell, they don't. Browsers implement the WHATWG's spec, which says not to do that and was created to supersede RFC 3986.
- bbayles 3y agoWHATWG rejects ".://", yeah? It's not the most readable spec, but there's a tester here: https://jsdom.github.io/whatwg-url/ https://jsdom.github.io/whatwg-url/ I recently published bindings for ada (an implementation of the WHATWG URL Spec) for Python with the hope of having something that follows a single standard.
- LegionMammal978 3y agoIndeed, ".://" is a hard error under the WHATWG URL spec. If the URL doesn't start with an ASCII alpha character, then the scheme start state transitions to the no scheme state [0]. In that state, if there's no base URL that the input is relative to, then parsing must fail [1]. However, "evil.com://good.com" is a valid URL string per WHATWG, since its state machine accepts "." within the scheme after the first codepoint. The resulting URL object has a scheme of "evil.com", a host of "good.com", an empty path, and a null port, query, and fragment. [0] https://url.spec.whatwg.org/#scheme-start-state https://url.spec.whatwg.org/#scheme-start-state [1] https://url.spec.whatwg.org/#no-scheme-state https://url.spec.whatwg.org/#no-scheme-state
- chrismorgan 3y agoIt’s not fair to call it a hard error: it’s only invalid as an absolute URL. As a relative URL, it’s fine, just like “example.com” is invalid as an absolute URL but valid as a relative URL.
- LegionMammal978 3y agoTrue; I neglected to mention relative URL parsing, mostly since most URL manipulation I've personally done has been with absolute URLs.