7 ms·
You know I was actually really curious about this so I went back to the HTML and URL W3C standards and surprisingly they don't actually have any definitions of
by jedimastert 5mo ago
You know I was actually really curious about this so I went back to the HTML and URL W3C standards and surprisingly they don't actually have any definitions of format other than being percent encoded. One might conflate query strings with "form-urlencoded"[0] query strings, which is one potential interoperability format, but in general a queries string is just any percent encoded string following a "?" in a url[1], and just another property in the "URL" HTML object that can be used in the generation of a response. While additionally there is a URLSearchParams object that is the result of parsing the query string with the form-urlencoded parser, this is simply an interoperability layer for JavaScript.
I'm going to be honest, I was pretty geared up to have a contrarian opinion until I looked at the standards but they're actually pretty clear, a 404 could be a proper response to unexpected query string; query string is as much part of the URL API as the path is and I think pretty much everyone can acknowledge that just tacking random stuff onto the path would be ill advised and undefined behavior.
[0]: https://url.spec.whatwg.org/#application/x-www-form-urlencoded https://url.spec.whatwg.org/#application/x-www-form-urlencod...
[1]: https://url.spec.whatwg.org/#url-class https://url.spec.whatwg.org/#url-class
- nrds 5mo agoWait until you realize that the difference between path and query string is entirely arbitrary and decided by the server. Query strings should never have existed. They are an implementation detail of CGI webservers that leaked all over everything and now smells really bad.
- jolmg 5mo agoIt's arbitrary to a degree like the difference between using an attribute or child element in XML, but it's not entirely arbitrary. If you want to include data in the URL that's not part of the hierarchy of the path, query strings are good for that.
- gpvos 5mo agoQuery strings existed before CGI did, and the way they're defined to be filled in from web forms is quite useful; I wouldn't want to need Javascript to fit that into path format. There's nothing wrong about having things decided by the server; I don't get that part of your argument at all.
- cobbzilla 5mo agoMaybe dumb question: how does the server “decide” anything other than what file to serve? Today we have many choices but back in the day CGI was the first standard way to do it. So yes query parameters existed before CGI but to use them you had to hack your server to do something with them (iirc NCSA web servers had some magic hacks for queries). CGI drove standardization.
- stirfish 5mo agofunc specialHandler(w http.ResponseWriter, r *http.Request) { if time.Now().Weekday() == time.Tuesday { http.NotFound(w, r) return } fmt.Fprintln(w, "server made a decision") } Your server can make decisions however you program it to, you know? It's just software. Forgive the phone-posting.
- deleted 5mo ago[deleted]
- cobbzilla 5mo agoand what server software is running this code in 1995?
- lispwitch 5mo agoCL-HTTP or AOLserver
- cobbzilla 5mo agosure looks like VB there, what’s the plugin? Didn’t see anything like that before.
- heavensteeth 5mo agoThat's Go.
- mikeocool 5mo agoI dunno, it seems like the fact that we arrived at a fairly standard structure for URL paths that works pretty well is not a bad outcome. Seems a lot better than the other potential world we could lived in, where paths were a black box and every web server/framework invented their own structure for them.
- hamburglar 5mo agoMy next website is going to have the path portion of the URL be a base64 encoded ASN.1 blob.
- chrismorgan 5mo agoSo long as it starts with a slash, go ahead! See how long it takes for someone to figure it out. It’s your website. Have fun with it! Do dumb things! :-)
- rkeene2 5mo agoMake sure you use URL-safe base64 or the portions that looks like a path can get mangled MII//epi Is converted to MII/epi
- yencabulator 5mo agoThat would be broken software. https://en.wikipedia.org/wiki/// https://en.wikipedia.org/wiki///
- gritzko 5mo agoIn my current project I use URIs to refer to absolutely any entity in a git(-ish) repo. Files, branches, revisions, diffs, anything. URI turns out to be a really good addressing scheme for everything. Surprise. But the most used and abused element is always the path. Query takes a lot of that mess away. Might have been unmanageable otherwise. https://github.com/gritzko/beagle https://github.com/gritzko/beagle
- 5mo ago
- paulddraper 5mo agoHow do you figure? Paths are hierarchical; query strings are name/value. (Note I speak of common usage.) You can create a different convention, but that one is pretty dang useful.
- halayli 5mo agoNothing you said here is correct. Paths, query strings, and fragments are all well defined entities. https://datatracker.ietf.org/doc/html/rfc3986#section-3.3 https://datatracker.ietf.org/doc/html/rfc3986#section-3.3
- sroussey 5mo agoIt’s a string between ? and # isn’t well defined. Or it is and it says very little.
- pverheggen 5mo agoNot entirely arbitrary - forms that use the GET method instead of POST will append form values as query params. For sites without Javascript, it's great for things like search boxes, tables with sorting/filtering, etc. instead of POST, since it preserves your query in the URL. https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/form#method https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/...
- msandford 5mo agoIt has always amazed me how much trouble the SPA folks are willing to go to in order to slowly rebuild just normal boring URLs with querystrings because users demand deep linking and back buttons and the like. Or you could accept that you're probably going to need a round trip to the server and use a normal URL and it's fine. For all but the absolute biggest websites in the world, anyhow. At Facebook or Google scale yeah it's needed.
- wongarsu 5mo agoBack in the day it was reasonably common for CMSs and forums to only have an index.php, and routing entirely by query string (in form-urlencoded form, people were not savages). So you would have index.php?p=home and index.php?p=shop. Or index.php?action=showthread&forum=42&thread=17976. It should be immediately obvious that in that scheme 404 is indeed the correct answer to unknown query parameters In fact lots of sites still work like that, they just hide it behind a couple rewrite rules in apache/nginx for SEO reasons
- Semiapies 5mo agoIf you're routing like it's 1999, sure, 404. On the other hand, if it's a CRUD app and you're filtering a list of entities by various field values? Returning that no items matched your selection (or an empty list, if an API) makes more sense than a 404, which would more appropriate for an attempt to pull up a nonexistent entity URI.
- Sander_Marechal 5mo agoThere is no reason you can return that "no items matched your selection" with a 404 HTTP response code instead of a 200.
- threatofrain 5mo agoA response code of 204 seems more appropriate but the problem is you're not allowed to send further information, which would make that descriptive response... not descriptive enough.
- IgorPartola 5mo agoI think of it like this: /users/ returns a 404 in an API means that this resource does not exist. As in, this is not a part of the API. /users/123 returns a 404 means this user record does not exist. Yes this means that a 404 is context dependent but in a way that makes it easier for a human to think of and reason about.
- qiller 5mo agoInterestingly, quite a few places that should treat query strings transparently make a lot of assumptions about their structure. We ran into that when picking a new CDN, some providers didn't handle repeat parameters (?a=1&a=2) correctly.
- sroussey 5mo agoWhat’s do you mean by correctly?
- kstrauser 5mo agoIncorrectly would be processing the query string and deduping keys. Correctly would be passing it through as-is, or at least only lightly processing it, like normalizing escaping or such.
- sroussey 5mo agoIndeed I would expect pass through with no changes. Though there are “smart” CDNs that will resize images etc. all beats are off for those.
- jedimastert 5mo agoFor anyone curious like I was, form-urlencoded and the URLSearchParam API says that params should not be deduplicated or reordered. "Get" will get the first value with the given name, and GetAll will get a list of all values https://url.spec.whatwg.org/#dom-urlsearchparams-get https://url.spec.whatwg.org/#dom-urlsearchparams-get
- ompogUe 5mo agoSomething I discovered looking back at some old sites: "pages" defined by URL params don't always make it into the Wayback Machine.
- bawolff 5mo ago> I was pretty geared up to have a contrarian opinion until I looked at the standards but they're actually pretty clear, a 404 could be a proper response to unexpected query string; query string is as much part of the URL API as the path is and I think pretty much everyone can acknowledge that just tacking random stuff onto the path would be ill advised and undefined behavior. This feels like a technically correct is the best kind of correct situation. Like technically, yeah web servers may respond 404 if they dont understand a query parameter, but in practise that is not how urls are conceptualized normally.
- TZubiri 5mo agoWhatwg is for html, try the IEEE http rfcs
- jedimastert 5mo agoThe IEEE rfcs does define a spot for the query string, but doesn't really say what to do with it. https://datatracker.ietf.org/doc/html/rfc3986#section-3.4 https://datatracker.ietf.org/doc/html/rfc3986#section-3.4
- TZubiri 5mo agoTry 1866 and 1867 Html is relevant historically as that syntax comes from forms. It was historically a sort of API between the browser client and the server, so yeah. But it's pretty well defined since 1995
- nofriend 5mo agoStandards are just commonly accepted behaviour that somebody chose to write down somewhere. There are a great number of commonly accepted behaviours that nobody's ever bothered to encode into a formal standard, but where failure to follow the accepted practice will result in widespread breakage. There are also a great many "standards" that you would be a fool to follow to the letter. In the OP case, the only thing that will break is people trying to visit their site, who will presumably simply press the back button on their browser and go about their day. They can decide for themselves if that is an acceptable casualty. But it isn't definitionally acceptable because no standard says it isn't (nor would is suddenly become unacceptable because a standard said it was...)
- chrismorgan 5mo agoYeah, URLs really don’t have much in the way of semantics. Path is clearly intended for hierarchical data and query for non-hierarchical data, and there are strong customs, some commonly supported or even enforced by libraries, but no actual rules. Ultimately, it’s just a string that the server can decide what to do with. The really funny thing about this is that, when I was worrying about possible side effects if I responded 404, I somehow completely forgot how much of the web’s history the path has been useless for. Paths have won. No one really starts new things with URLs like /item?id=… any more. Yay!
- fpoling 5mo agoWikipedia web server treats anything after /wiki/ literally as the name of the article. So en.wikipedia.org/wiki/// is the article about C++ style comments
- chrismorgan 5mo agoOh, magnificent. Lovely high-profile example to add about empty path segments being meaningful.
- chii 5mo agoi wonder if it ought to be `/wiki/%2F%2F` instead...
- morganastra 5mo agoit looks like it goes to a disambiguation page for what "//" could refer to now (C++ style comments being the top entry), but that's delightful!
- dylan604 5mo agoWouldn't a generic 400 be better. It's not that the page wasn't found, but you've sent something that was not an accepted request. Fix your request and try again is how I've read it, and that's how I use it in the APIs I provide. I prefer it over 406 since it's not my end that can't process it. If your query string is tacking extra stuff trying to break things or just because your request wasn't crafted per the docs, then it's on you.
- Ekaros 5mo ago406 would be wrong for me. As it is to be used when client sends Accept: header and server cannot fulfil that. HTTP return codes get quite specific when you read the actual description and not just name.
- socalgal2 5mo agoThe No-Vary-Search (proposal?) https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/No-Vary-Search https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/... effectively lets you specify what parts of a query are relevant. So for example url?a=b&c=d matches url?c=d&a=b in terms of caching