8 ms·
URLs: It's Complicated
- leifg 5y agoIt seems like the colon is too ambiguous (is used as a protocol delimiter, delimiter for user/pass, delimiter for port). Reminds a little bit of Java labels where you can do this: public class Labels { public static void main(String args[]){ https://hn.ycombinator.com for(int i=0; i<10; i++){ System.out.println("......."+i ); } } } the https: is a label named https and everything after the colon is a comment so this is valid code.
- jackewiehose 5y ago> It seems like the colon is too ambiguous (is used as a protocol delimiter, delimiter for user/pass, delimiter for port). and because that was still too boring they came up with ipv6
- deleted 5y ago[deleted]
- mananaysiempre 5y agoIIUC the IPv6 weirdness here is simply due to very unfortunate timing: IPv6 was being finalized at a time (first half of the 90s) when the Web (and with it URLs) was already nearly frozen but still not obviously important.
- chungy 5y agoThe colons also make IPv6 addresses unambiguous to IPv4 notation.
- gpvos 5y agoNot in URLs, but related: Larry's 1st Law of Language Redesign: Everyone wants the colon Larry's 2nd Law of Language Redesign: Larry gets the colon https://thelackthereof.org/Perl6_Colons https://thelackthereof.org/Perl6_Colons
- prepend 5y agoIt doesn’t seem complicated at all. Complicated to me means difficult to understand. This just involves reading the spec and it all seems pretty simple and consistent. Complicated doesn’t mean “new to me.” If I haven’t read a man page, that doesn’t mean the command is complicated.
- treve 5y agoI'm trying to figure out what your point is. Is it a criticism of the article? Are you sharing with us that you are clever? I'm trying to give it a generous interpretation, but I'm having a hard time. So it's not complicated to you... what made you want to share this?
- Stratoscope 5y ago> This just involves reading the spec and it all seems pretty simple and consistent. Worth noting from the article: > And this is one of the main take-aways here: while the URL specification prescribes or allows one thing, different clients and servers behave differently.
- kube-system 5y agoIf we're being sticklers for reading the docs... > com·pli·cat·ed | ˈkämpləˌkādəd | > adjective > 1 consisting of many interconnecting parts or elements
- deleted 5y ago[deleted]
- alphabet9000 5y agoEven browser developers have made mistakes as a result of the complexity of the spec, resulting in things like CVE-2018-6128 [0] happening. [0] https://bugs.chromium.org/p/chromium/issues/detail?id=841105 https://bugs.chromium.org/p/chromium/issues/detail?id=841105
- sitdown 5y agoLayouts using <table>s are complicated too. For example, this page has a ~7800px-wide <pre> tag in a <table> that's 720px wide.
- scandinavian 5y agoSpecifically using another font for the code tag then the rest of the blog to hide the difference between ⁄⁄ and // seems weird. I get that it wouldn't be interesting if not doing that, but doesn't that just show that it's really not as complicated as you make it out to be?
- teknopaul 5y agoURLs are not complicated, unless you complicate them. foo|foo -foo 's^foo^foo^'"">foo 2>>foo is not a very good example for teaching the structure of the the command line. Pick a better one. It's simple.
- LambdaComplex 5y ago"The average URL" and "what is allowed by the URL specifications" are two very different things. (And the same could be said about your command line example)
- zepearl 5y agoAll extremely useful: the overview, the examples and the comments. A few months ago while writing a bot/crawler I searched for hours for something like this, but I found only full specs or just bits and pieces scattered around that used different terminology and/or had different opinions. In the end I didn't even clearly understand what should be the max total URL length (e.g. mixed opinions here https://stackoverflow.com/questions/417142/what-is-the-maximum-length-of-a-url-in-different-browsers https://stackoverflow.com/questions/417142/what-is-the-maxim... - come on, a xGiB long URL?) => most of the time 2000 bytes is mentioned but it's not 100% clear. Writing a bot made me understand 1) why browsers are so complicated and 2) that the Internet is a mess (e.g. once I even found a page that used multiple character encodings...). My personal opinion is that everything is too lax. Browsers try to be the best ones by implementing workarounds for stuff that does not have (yet) or does not comply to a spec => this way it can only end up in a mess. A simple example is the HTTP-header "Content-Encoding" ( https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Content-Encoding https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Co... ) which I think should only indicate what kind of compression is being used, but I keep seeing in there stuff like "utf8"/"image/jpeg"/"base64"/"8bit"/"none"/"binary"/etc... and all those pages/files work perfectly in the browsers even if with those values they should actually be rejected... .
- mananaysiempre 5y agoThe use of Content-Encoding for compression is actually something of a historical wart: what was intended to be used for that purpose is Transfer-Encoding, but modern browsers don’t even send the TE header necessary to permit the HTTP server to use it (except for Transfer-Encoding: chunked which every HTTP 1.1 client must accept), even though some servers are perfectly capable of it and all but the most broken will at least ignore it. Things like 7bit, 8bit, binary, or quoted-printable are not supposed to be in the HTTP Content-Encoding header, either, but their presence is at least somewhat understandable as they are valid in the MIME Content-Transfer-Encoding header, and HTTP originally shares much of its infrastructure with MIME (think Content-Disposition: attachment). I guess what I’m getting at here is that the blame for the C-E weirdness lies in large part on the browsers, which could’ve made a clean break and improved the semantics at the same time by using T-E, but instead chose to initiate a chicken-and-egg dilemma out of a desire to support broken HTTP servers from the last century. (The intended semantics is that C-E, an “end-to-end” header, says “this resource genuinely exists in this encoded form”, while T-E, a “hop-to-hop” header, says “the origin or proxy server you’re using incidentally chose to encode this resource in this form”; this is why sometimes the wrong combination of hacks in the HTTP server and the Web browser will lead you to downloading a tar file when you expected a tar.gz file.) The use of “gzip” as the compression is also a wart, because it’s “deflate” (which is what you want: DEFLATE compression with a checksum) with a useless decompressed filename (wat?) + decompressed mtime (double wat?) header stacked on top.
- mananaysiempre 5y agoJust to share a little more of the weirdness (discovered while reading a couple of the historical URL & URI RFCs several days ago): Per the original spec, in FTP URLs, - ftp://example.net/foo/bar will get you bar inside the foo directory inside the default directory of the FTP server at example.net (i.e. CWD foo, RETR bar); - ftp://example.net//foo/bar will get you bar inside the foo directory inside the empty string directory inside the default directory of the FTP server at example.net (i.e. CWD, CWD foo, RETR bar; what do FTP servers even do with this?); - and it’s ftp://example.net/%2Ffoo/bar that you must use if you want bar inside the foo directory inside the root directory of the FTP server at example.net (i.e. CWD /foo, RETR bar; %2F being the result of percent-encoding a slash character).
- jfrunyon 5y ago> what do FTP servers even do with this? Pretty sure CWD by itself isn't even valid (at least RFC959 assumes it has an argument), and therefore // isn't valid in FTP URLs. The %2Ffoo/bar is needed because of the fact that FTP CWD and RETR paths are system dependent (with, theoretically, system dependent path separators), but URLs are not, so the FTP client breaks the URL on / and sequentially executes CWD down the tree so that it doesn't need to know what it's connected to. In other words: URL paths are not system paths, and it's a mistake to think of them as such. (Alternate in other words: FTP is awful)
- mananaysiempre 5y ago> Pretty sure CWD by itself isn’t even valid [...] and therefore // isn’t valid in FTP URLs. So, I looked it up carefully and it appears that (despite the promises in later RFCs such as 2396 and 3986) the current specification of the ftp scheme is still the ancient RFC 1738 which predates not only the URL / URI distinction but even the notion of relative URLs. In §3.2.2 <https://tools.ietf.org/html/rfc1738#section-3.2.2 https://tools.ietf.org/html/rfc1738#section-3.2.2> it specifically says that a null segment in the path should result in a “CWD ” command (i.e. CWD, space, null string argument) being sent to the FTP server, going against both the current RFC 959 and its predecessor 765 (apparently the earliest formal specification of FTP to include CWD) which require the argument to CWD to be non-null. Thus apparently a conformant implementation of the ftp URL scheme cannot be a conformant implementation of an FTP client. Joy. It still seems unlikely that Berners-Lee et al. would specifically call this case out if it were useless at the time... What were the servers that made this necessary, I wonder? > FTP CWD and RETR paths are system dependent (with, theoretically, system dependent path separators), but URLs are not Thank you, that’s the insight that I was missing. So a %2F inside an ftp URL component is just performing a (sanctioned) injection of the (supposedly UNIXy) server path syntax. > FTP is awful I’d go with “unbelievably ancient, with the attendant problems”, but yes. Funny how it still manages to be better than everything else (that I know) at transferring files by not multiplexing control and data onto the same TCP connection. (I think HTTP over QUIC can do this as well?)
- surfingdino 5y agoI have come across even more issues caused by IRIs used incorrectly in place of URIs by a popular web framework, causing havoc with OAuth redirects. https://en.wikipedia.org/wiki/Internationalized_Resource_Identifier https://en.wikipedia.org/wiki/Internationalized_Resource_Ide...
- jfrunyon 5y ago> making this is a valid URL: https://!$%:)(*&^@www.netmeister.org/blog/urls.html https://!$%:)(*&^@www.netmeister.org/blog/urls.html Uh, no. "%:)" is not <"%" HEXDIG HEXDIG> nor is % allowed outside of that. (Although your browser will likely accept it) > This includes spaces, and the following two URLs lead to the same file located in a directory that's named " ": > https://www.netmeister.org/blog/urls/ https://www.netmeister.org/blog/urls/ /f > https://www.netmeister.org/blog/urls/%20/f https://www.netmeister.org/blog/urls/%20/f > Your client may automatically percent-encode the space, but e.g., curl(1) lets you send the raw space: Uh, no. Just because one of your clients is wrong and some servers allow it doesn't mean it's allowed by the spec. In fact, the HTTP/1.1 RFC defers to RFC2396 for the meaning of <abs_path>: <path_segments> which begin with a /. What is <path_segments>? A bunch of slash-delimited <segment>s. What is <segment>? A bunch of <pchar> and maybe a semicolon. What is <pchar>? <unreserved>, <escaped>, or some special characters (not including space). What is <unreserved>? Letters, digits, and some special characters (not including space). What is <escaped>? <"%" hex hex>. Most HTTP clients and servers are pretty forgiving about what they accept, because other people do broken stuff, like sending them literal spaces. But that doesn't mean it's "allowed", that doesn't mean every server allows it, and that doesn't mean it's a good idea. > That is, if your web server supports (and has enabled) user directories, and you submit a request for "~username": [it does stuff] Uh, no. If you're using Apache, that might be true. As you mentioned, this is implementation-defined (as are all pathnames). > Now with all of this long discussion, let's go back to that silly URL from above: ... Now this really looks like the Buffalo buffalo equivalent of a URL. Not really. > Now we start to play silly tricks: "⁄ ⁄www.netmeister.org" uses the fraction slash characters You are aware that URLs predate Unicode, right? Not to mention that Unicode lookalike characters are a Unicode (or UI) problem, not a URL problem? > The next "https" now is the hostname component of the authority: a partially qualified hostname, that relies on /etc/hosts containing an entry pointing https to the right IP address. Or on a search domain (which could be configured locally, or through GPO on Windows, or through DHCP!). Or maybe your resolver has a local zone for it. Or maybe ...