8 ms·
URL query parameters and how laxness creates de facto requirements on the web
- Supermancho 6y ago> If a HTTP request has unexpected and unsupported query parameters, such a GET request will normally fail. When I made this decision it seemed the cautious and conservative approach, but this caution has turned out to be a mistake on the modern web. For a web page it might be a mistake. For an API it's often not a mistake, depending on the utility of the API. eg analytics reporting.
- evmar 6y agoMy favorite story to tell in this area: many years ago in the early days of Chrome, there was some logic that followed HTTP redirects that had a cap on how many redirects to follow before giving up. (The rationale here is that if you hit too many in a row you probably had encountered a broken site , like one where /foo?q=1 redirects to /foo?q=2 redirects to q=3 and so on). As I recall, the cap on redirects was 10, and then Darin (who had previously worked on Firefox) was like "no, you have to allow up to 30 or you break the New York Times". And so, it was upped to to 30. Intuitively I'd think a site that redirected you 10 times was broken, but I guess someone built one that needed more. PS: Looking online now to confirm my facts, I see one claim that both Chrome and FF limit redirects to 20. So either my memory exaggerated it as 30 or they both managed to lower the limit at some point.
- onion2k 6y agoI kind of wish the cap had remained at 10, and the redirection on the NYT was left as a problem for them to fix...
- jrockway 6y agoUpstart browsers are not in the position to tell websites what to do; they will just throw up a "your browser is not supported" banner and forget about you. That was where Chrome was in its early days (you were supposed to be using IE 6), and that's where Firefox is now. We never let economics solve this problem, because nobody pays for web browsers with money. If there was a $10 "sorry, The New York Times violates standards so we don't support it" and a $100 "we'll work around any bug large or small to make it work for our customer" browser sites would be spending money to let the $10-browser-owners see their ads. But, every browser is run as a charity, so nobody gets to pay the actual costs of the workarounds (except maybe in terms of more security vulnerabilities, higher RAM usage, etc.) I'm not saying this is good or bad, it's just how it is. Browsers are free and will be blacklisted by sites that don't like how they behave. So the incentive for browsers is to do what the sites want, rather than to strictly conform to standards. Shrug.
- rmdashrfstar 6y agoWhat could they be doing with all those redirects?
- p_l 6y agoRun you through all the spying services' cookies & other fingerprinters. Just like Twitter's t.co.
- SquareWheel 6y agoThe first two are often: nytimes.com > https://nytimes.com https://nytimes.com > https://www.nytimes.com And from there, who knows. Some websites append a trailing slash; sometimes they removed it. Some websites send you to /index.php or /home.aspx rather than sitting in the root. Subpages may use a redirect to update a slug if there's been a title change. Or maybe the website just moved to a new format altogether and is updating old links. I have seen examples of 4 or 5 legitimate redirects to send users to the correct place. I'm having trouble coming up with 10 possible reasons, though.
- Vinnl 6y agoHa, that must be for academic articles as well. They all (well, almost all) have an ID that you can look up at doi.org, which is supposedly kept up-to-date and links to the article at the publisher's website, wherever that is. But then in practice, that publisher's website adds another bunch of redirects because why not.
- drchickensalad 6y agoI had an edge case with an authentication server (think okta) wrapped around a school login system. In occasional cases, with certain clients of the server doing a couple redirects of their own, we'd hit that 20 cap. It's not any individual system being irresponsible, just reusing other systems as they're meant to be. It's kind of like saying having a call stack depth of 50 in software is never acceptable.
- paledot 6y agoWhile I respect sticking to your guns, I don't think we're putting that genie back in its bottle. Maybe a reasonable middle ground, like a less-fraught version of the HTTPS redirect, would be to have any requests containing unknown query parameters redirect to the URLs with those parameters stripped. It does nothing to deter the people slapping them on there, but at least your visitors don't spread that plague by sharing your URL. Of course, it all comes with the problem of needing to define up-front which query parameters your app accepts. Easy in a small app, wildly impractical in a large one with plenty of legacy code.
- klodolph 6y agoThat's exactly what the 301 redirect is supposed to do. "Moved permanently."
- evilotto 6y agoit also breaks the ad and click tracking that those parameters are for in the first place. You're probably happy to do that, and that's just fine.
- twic 6y agoBetter yet, send people to a trampoline page saying "You have followed a broken link. Did you mean to visit ...". Let people get the content, but rub their face in the fact that Facebook is breaking things.
- Joker_vD 6y agoWhat is that even supposed to accomplish other than signalling to the visitors that a) you dislike Facebook, b) you don't mind punishing others for what you perceive to be Facebook's wrongdoings? Make them abandon Facebook? Make them make Facebook to do things your way?
- twic 6y agoIt probably wouldn't have any effect on Facebook. But if every link from Slack was hitting a page like this, Slack would fix it, because it makes them look stupid to their technical audience.
- gruez 6y ago>Today's new query parameter is 's=NN', for various values of NN like '04' and '09'. I'm not sure what's generating these URLs, but it may be Slack For a vague parameter like "s", my bet it's probably due to bad code rather than intentionally added. eg. someone constructs a URL but forgets to escape the components, and later (or downstream) some other programmer decides to "fix" it by using encodeuri (https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/encodeURI https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...), which "works", but mangles the url
- pierregoutagny 6y agoActually, I think Twitter adds "s=NN" to their URLs when you share them depending on your device. It's something like NN=19 for Android, 20 for Windows, and 21 for iOS.
- Akronymus 6y agoI REALLY hate that sharing URL. Makes me have to do another click to see the actual thread with replies directly. A subtle, but, to me, important difference.
- gpvos 6y ago"s" is a really dangerous parameter to use, though. 100% guaranteed that lots of sites already use it. I mean, fbclid isn't fun, but at least it's fairly unique.
- inopinatus 6y agoWell, I believe in Postel’s Law. Any opinions on the utility or danger of redirecting to the canonical URL instead? E.g. I don’t want anyone’s campaign tracking tokens in the referer of our outbound links, or in URLs being shared around by copy-paste from the address bar. However I also don’t want to inadvertently create a mechanism for someone to mechanically enumerate our routes (which we already see routinely attempted via return-path parameters), or walk into any similar trap besides.
- ryan-c 6y agoI think this might be a good use case for history.replaceState - it'll change the URL in the address bar without actually redirecting.
- techbio 6y agoReferenced (only partially quoted, and wholly unheeded) in the article, Postel's Law, also known as the Robustness Principle, says "be liberal in what you accept, and conservative in what you send." https://en.m.wikipedia.org/wiki/Robustness_principle https://en.m.wikipedia.org/wiki/Robustness_principle
- oaiey 6y agoBypass the (proxy) cache: add a query parameter with a random value in it. Forgot, why I needed it but back in the day it was a big deal.
- chii 6y agothe old jquery would automatically add this parameter to your get requests to bust the cache.
- panopticon 6y agoI worked at a few marketing shops a while ago where we would manually increment a cache-busting query parameter any time we changed static assets (e.g., `/css/style.css?v=20100908`). We configured Apache to tell the client to cache these resources for a year, and this was a fairly common way of getting around that. Forgetting to update this ended up be a common source of bugs.
- wereHamster 6y ago> One of the ways that DWiki (the code behind Wandering Thoughts) is unusual is that it strictly validates the query parameters it receives on URLs, including on HTTP GET requests for ordinary pages. If a HTTP request has unexpected and unsupported query parameters, such a GET request will normally fail. This goes counter to the robustness principle: Be conservative in what you do, be liberal in what you accept from others (https://en.wikipedia.org/wiki/Robustness_principle https://en.wikipedia.org/wiki/Robustness_principle).
- olliej 6y agoWhich we’ve increasingly moved away from because of some of the absurdities (and security problems) that it’s caused. Nowadays the w3c and es committees work hard to try and ensure no ambiguity is left in the spec, and have tests available early on, and do interop comparisons between engines to see if there’s anything missed.
- im3w1l 6y agoA big issue with being accepting is that you risk breaking things tomorrow when you no longer want to be as accepting. And it can prevent adding features that would play badly with others' out of spec implementations. Like if twitter adds these s=NN parameters, that means you can't use an s=NN parameter for your own purposes.
- pwdisswordfish4 6y agoAll the worse for the ‘robustness’ principle then. https://tools.ietf.org/id/draft-iab-protocol-maintenance-04.html https://tools.ietf.org/id/draft-iab-protocol-maintenance-04.... (Shame this draft just expired before being finalised.)
- treve 6y agoThe robustness principle is widely considered to be a bad practice these days. Strictly failing is a good thing, generally. Accepting arbitrary query parameters is pretty much a de facto exception though
- lixtra 6y agoThe enpoint is possibly not the only adressee of the parameter. They may go to some middleware component. Or the JavaScript in the browser. So it makes sense to take parameters as messages that you ignore when unknown.
- dgoldstein0 6y agobut then the middleware or js could just strip out it's extra params. Though that invites it's own problems - e.g. if the application actually used the `s` param, and something came along and started using it too... but this case is in general a world of hurt. I'm not really sure why we'd want a world where the server doesn't know about all query args, but seems like we're already stuck with this.
- Animats 6y agoDo those extra parameters break caches? The browser's cache doesn't know it's the same URL. Does Cloudflare? Akamai?
- EE84M3i 6y agoYes and no. I believe all CDN have config available to put no query params, all query params or specific query params into cache keys. Whether people us that options is a different question.
- dehrmann 6y agoIt's worth mentioning Postel's Law: https://en.wikipedia.org/wiki/Robustness_principle https://en.wikipedia.org/wiki/Robustness_principle That said, query parameters on a GET are one thing, but unexpected arguments on CLIs are another, especially when the command can change/delete files, so I can see both sides of this.
- Thorrez 6y agoSomething similar to URL query parameters is cookies. Most servers will ignore unknown cookies, but not always. The Great Suspender Chrome extension inserted tons and tons of cookies into various domains for some users, and there were so many cookies that this broke Google sites for some users. https://github.com/greatsuspender/thegreatsuspender/issues/537 https://github.com/greatsuspender/thegreatsuspender/issues/5...
- inoffensivename 6y agoYears ago I worked on Google Maps, we spent an inordinate amount of time making sure that we didn't break backwards compatibility. Sometimes we would be trying to refactor some old code, but we'd get stuck trying to support some ancient client that was sending us like 4 requests a day with some wacky query parameters. On the other hand, our Maps still worked on all 5 Blackberries that people were still using.
- jooize 6y agoWere there different user agents or did all clients have to use the same interface?
- mattmanser 6y agoErr, from a user perspective you actually spent an inordinate amount of time changing things for the sake of it. It was really annoying. Google APIs have always been some of the worst. It honestly felt like ever time I had to use the maps API for a client, you'd changed the whole thing.
- inoffensivename 6y agoYeah, I think people felt less compunction changing the published APIs. I was talking about the "internal" APIs, especially the tile API which still serves ancient clients.
- disgruntledphd2 6y agoThat is an extremely strange order of priorities, right?
- mschuster91 6y agoProbably because the "internal" API was used by someone who paid a lot of dollars?
- hyperdimension 6y ago
- wodenokoto 6y agoReading this article I start to wonder about how common parameter collision is in the wild. Probably, your service doesn’t have any ‘utm_*’-parameters not intended but I could see ‘fbcid’ being used for other ids than just Facebook clients and ‘s’ is so generic that it must clash with thousands of services!
- sradman 6y agoPerhaps a general compromise is to generate a rel=canonical link [1] that only includes directly supported query parameters. [1] https://en.wikipedia.org/wiki/Canonical_link_element https://en.wikipedia.org/wiki/Canonical_link_element
- bandie91 6y agoIMO, query (search) parameters are originally designed for remote database search and forms. search criteria and form fields are often contain optional and conditional fields (conditional ie. only meaningful if some other field checked or having some specific value). but clients send them anyway because the logic (what is optional and which field depends on which) is known by the server. that's why client side programs consider server-side programs being lax accepting any parameter. however appending various cryptic 1-letter parameters without knowing the target http resource's semantics is unclever, since get parameters are part of the url, so changing them changes the url itself, so the new url does not necessarily points to the same resource. IMO, it's the web's originally visioned concept of identifier–resource-relation which is not adhered in today's websites. and this leads to dirty workarounds, unnecessary complexity and security holes.
- shp0ngle 6y agoI mean, that's almost the same debate as XHTML vs HTML5 of yesteryear... Some people want strict requirements, but "whatever works for the end-user" usually wins. ...and as a result, we got monstrosities like user agent strings. But that's something we need to deal with I guess
- thiht 6y agoI have a hard time understanding why one would care about such things. Just ignore additional query parameters you don't need and be done with it forever, what's the problem with that? Is it really worth writing a rant?
- ben509 6y agoThe reason is stated towards the end: > In general, any laxness in actual implementations of a system can create a similar spiral of de facto requirements. Those "de facto requirements" mean that writing something simple becomes inordinately complex because noone fixes broken clients. And you can't test changes because you don't have access to these broken clients. It also leads to nasty bugs. For example, if you ignore extra parameters and provide defaults for others, now a misspelled parameter will silently fail. The best example is how hard it is to write a browser. The standards are very complex, but it's magnified by the complexity of supporting "real," namely broken, HTML and CSS. That's contributing to the browser mono-culture.
- Joker_vD 6y agoWell, the alternative is that somehow you end with technology that no one bothers to use. Laxness allows to have multiple early, lazy implementations interoperate more or less correctly. Yes, in the later stages you end up with loads of technical depth and oligo- or mono-culture.
- alisonkisk 6y agoThe logic is incoherent. Strictness creates requirements too.
- dathanb82 6y agoThe article isn’t arguing against requirements existing, it’s arguing against being bound to de facto requirements because you allowed a client you don’t own to depend on behavior you didn’t specify. (See Hyrum’s Law: https://github.com/dwmkerr/hacker-laws#hyrums-law-the-law-of-implicit-interfaces https://github.com/dwmkerr/hacker-laws#hyrums-law-the-law-of...). By limiting the scope of unspecified behaviors, you also limit the scope of what makes for a de facto breaking change.
- cxr 6y agoGit is another offender here. When they decided to switch to the smart protocol by default, the Git people decided that the way to probe for server support was to append a query parameter to the resource. Extra parameters are ignored by default on Apache (or something), and it worked on whatever the committer tested it on, so that became how Git worked. Whereas the parameter is part of the smart protocol, any server that supports it will give a response in line with what the smart protocol dictates, and the client will proceed with having recognized a server with smart protocol support. Dumb servers who ignore extra query parameters and respond with the kind of thing that dumb servers respond with will be treated as having no support for the smart protocol, and the Git client will allow the connection to fall back to the older protocol. Meanwhile, you can have a server that only supports the dumb protocol but doesn't ignore extra parameters, and so the switch in the client defaults just broke all transactions with those servers. The Git developers either didn't test it or didn't care. And something that makes the whole thing particularly silly/frustrating is that HTTP already has ways to do content negotiation which could have been used in lieu of the boneheaded scheme they came up with.