34 ms·
When the robustness principle meets the tragedy of the commons, good people get driven to the brink of madness. There are surely arguments for allowing any old
by tfm 10y ago
When the robustness principle meets the tragedy of the commons, good people get driven to the brink of madness.
There are surely arguments for allowing any old garbage in URLs/URIs/IRIs - hey, it's easy and fun! - but, gee, it's not like bunging an identifier through a URI escape function is going to triple the code base.
Compare that with the surface area for potential bugs in the parsing code which somehow didn't account for several megabytes of whitespace embedded in a URL or multiple code pages in a single string or whatever zalgoesque horror is showing up on the security bulletins this week. It's a lot easier to proof code that only needs to deal with a strict subset of 7-bit clean ASCII and can politely decline embedded emojis.
When one web client starts being overly zealous in what it accepts, that puts an implicit onus on everybody else to start accepting that too ("all browsers do"!). Where would you draw the line? Well, we've got a couple of RFCs lying around, how about we go with those.
This stuff doesn't have to be hard. Surely the act of issuing an HTTP redirect isn't the Last Great Unsolved Problem of web engineering! I say rejoice in the beauty of URIs in canonical form. Be miserly in what you accept for a change.
- tuukkah 10y agoMost of the world is outside the (English-speaking) United States of America and doesn't speak an ASCII language. That's why Unicode exists - couldn't we all just use it? That's what the IRI RFC says in essence. EDIT: Perhaps I misread and you're saying we need two layers like the RFCs intended it: the user-visible Unicode IRIs and the protocol-level ASCII URIs. cURL lives between these two and arguably would be simpler if there was no such separation and everything was Unicode to begin with.
- tfm 10y agoAgreed. I love me all the Unicodes, all the time! For typing, for display, for whatever internal app logic, variable names, the names of children etc. The IRI RFC does however specify that for transmission/interchange the IRIs need to be percent-encoded to form valid URIs. That's pretty straightforward! A good user agent will fully transcode whatever characters I type in the address bar, generating a valid IRI. A bad user agent will complain and say that whatever it was I typed was very nice but not a valid interweb, and I have a bad day. A broken user agent may generate some broken encoding (or just pass on whatever I typed in UTF-8 or CP-1252), and some web admins may have a bad day, if they weren't already. But the location bar is not the only place that URLs come from (less and less every day). When a web application (e.g. 301 redirect) or a resource (e.g. external document reference) uses an RFC-3987 noncompliant identifier, just how far should the user-agent bend over backwards to fetch the resource? Some browsers seem to do a lot, but it's not clearly documented just how much (see elsewhere in the comments), and among the article's laments was the observation that when any browser accepts dodgy identifiers, it places a burden on everyone else to accept the same. Seems that WHATWG hopes to formally codify just how dodgy you can get away with without the identifier being "too dodgy", but I get the impression that this will be an expansive definition ("hey, if they did it, you can do it! we believe in you"). This is an understandable approach, as they are trying to document the state of the web, but it does put extra burden onto anyone dealing methodically with URLs henceforth, rather than being strict and drawing some clear boundaries up for people constructing URLs. I'm expecting to see a lot of "SHOULD" and not a lot of "MUST NOT" :-/
- LoSboccacc 10y agoSure! Which unicode? Utf-8? -16? WTF-8 to allow for robustness in parsing?