10 ms·
I’ve banned query strings
Related: https://susam.net/no-query-strings.html https://susam.net/no-query-strings.html
- moritzwarhier 5mo agoThis is cool and creative! It uses 4xx, but not just 400 :) https://chrismorgan.info/no-query-strings?why=unknown https://chrismorgan.info/no-query-strings?why=unknown
- gtowey 5mo ago"wander console" sounds like they're just web rings re-invented. In the era of forced feeds by giant corporations which consist of the things they want you to see, I've wondered if this old idea would make a comeback. Human curated content from trusted people seems like the only way forward.
- SoftTalker 5mo agoFTA: It is also a bit like web rings except that the community network is not restricted to being a cycle; it is a graph and it is flexible.
- cosmicgadget 5mo agoIs it not a random walk? Might sound pedantic but if there is graph structure I am interested.
- susam 5mo agoI'll start with the clarification that the moderators have changed the URL of the original post from <https://susam.net/no-query-strings.html https://susam.net/no-query-strings.html> to <https://chrismorgan.info/no-query-strings https://chrismorgan.info/no-query-strings>. Hopefully, this will prevent any confusion about why we are discussing random walks in a post about query strings. Now let me answer your question. > Is it not a random walk? Might sound pedantic but if there is graph structure I am interested. The network is a directed graph. Every Wander Console declares a few other consoles as its neighbours. The person setting up the console decides who they want to list as their neighbours. So if we call the network graph X, then the set of vertices is: V(X) = the set of all URLs that point to Wander Consoles and the set of directed edges is: E(X) = {(u, v) in V(X) : u declares v as its neighbour} The traversal between consoles is not strictly a random walk. If I could call it something, I would call it randomised graph exploration with frontier expansion. On each click of the 'Wander' button, the tool picks one console at random from the set of discovered consoles and visits that console. It then fetches the neighbours declared by that console and adds any newly discovered consoles to the set. The difference from a random walk is that the next console is not chosen from the neighbours of the last visited console. It is chosen from the whole set of consoles discovered so far. In other words, each click expands the known part of the graph, but the console used for that expansion is selected randomly from all discovered consoles, not just from the last console visited.
- cosmicgadget 5mo agoOh cool, I love this approach. Randomized graph traversal but ever visited node is a fast travel station. Great way to avoid running into dead ends.
- julianlam 5mo ago> After I implemented that feature, a page from one of my favourite websites refused to load in the console... the third URL returns an HTTP 404 error page. The website uses the query string to determine which one of its several font collections to show. Yes, let's unilaterally decide that query strings are bad because one website (ab)uses query strings to load different fonts. It's the query strings that are the problem, not the website! jfc. Look, I'm against utm fragments as much as the next guy, but let's not throw away a perfectly good thing because tracking is evil.
- ergonaught 5mo agoAdding your own garbage to someone else's URLs is in fact the problem. Could they handle your garbage better? Sure. Is your garbage still a problem? Yes.
- SoftTalker 5mo agoPostel's law worked OK when people operated in good faith. But today the internet is full of abusers. Rejecting requests that aren't exactly what they should be is probably the best policy now.
- wtallis 5mo agoPostel's law is typically stated as "be conservative in what you do, be liberal in what you accept from others". It's unfortunately common for people to ignore the first half and hallucinate a third clause demanding that the recipient stay silent about the errors they receive.
- jorams 5mo agoThe website uses the feature for its intended purpose. Adding random trash to the query string of another website assuming it'll ignore it is in fact a bad idea, always, even if you can usually get away with it.
- InsideOutSanta 5mo agoThat website is not abusing query strings, though, its usage of query strings is perfectly cromulent. And tfa is not saying not to use query strings, but not to append random garbage to other people's URLs.
- sigseg1v 5mo agoAdding query strings is one of those things that I think a lot of sites could get away with more easily if they were reasonable about it. A link that is "https:// web.site" is fine. A link that is "https:// web.site?via=another.site" is fine. A link that is "https:// web.site?fbm=avddjur5rdcbbdehy63edjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63edaaaddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednzzddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63ednddjur5rdcbbdehy63edn" is annoying as shit and I need to literally apologize to people after sending it if I forget to manually redact the query string. Don't abuse this.
- culi 5mo agoThere are addons to remove unnecessary params from the worst offending sites: https://www.google.com/search?q=clearurls+addon https://www.google.com/search?q=clearurls+addon
- franciscop 5mo agoThanks for removing the rest on that google link, the one I get after switching to "images" and back to "web" is this monstrosity: https://www.google.com/search?newwindow=1&sca_esv=8061bd9cb19cd450&sxsrf=ANbL-n7S60ZBdf0lh5kQ8RojJdQpnM0S5w:1778353180297&q=clearurls+addon&source=lnms&fbs=ADc_l-aN0CWEZBOHjofHoaMMDiKpeTF8ggB1qASWZfpybz5TQZmqMiWOgtbP_iLwZE3_BsqFrIkjQk30pNpcyOJjgYT1NYhSr_eVWusunSdIYLAa1WWhJm7VPvRsNUkHss5YZDSVhzEth7KnRsP0kwdL-3ylxxDz_j5WL-QtjJdzQePIWAeCwn7532w9WuSzSqnY0V2tn342eEk_wDwxk45MDY_JuA-5CA&sa=X&ved=2ahUKEwjH3uLs8ayUAxUghP0HHVXuOeIQ0pQJegQICxAB&biw=1296&bih=711&dpr=2.22 https://www.google.com/search?newwindow=1&sca_esv=8061bd9cb1... Edit: which luckily and sensibly Hacker News cuts short since it's 463 characters
- 1shooner 5mo ago>So I’ve decided to try a blanket ban for this site: no unauthorised query strings. His site returns (I think incorrectly) a 414 if a request includes a query string. If this protest is meant to advocate for the user, who presumably wasn't able to manage that string in the first place, why would you penalize them for it being there? Why not just use it as a cue to tell users how they can make this decision themselves (e.g. through browser tools)?
- bryanrasmussen 5mo agoIt's been years but I seem to remember there was a version of PLSQL server pages that would return 500 if you tried to pass in an unknown query string.
- jampekka 5mo ago"You could argue that I’m abusing 414 URI Too Long. I respond that it’s funnier this way. Other options I considered were: 400 Bad Request, the generic client error code, which is correct but boring; 402 Payment Required, and honestly if you want to pay me to make a particular URL with query string work, I’m open to it; 404 Not Found, but it’s too likely to have side effects, and it doesn’t convey the idea that the request was malformed, which is what I’m going for; and 303 See Other with no Location header, which is extremely uncommon these days but legitimate. Or at least it was in RFC 2616 (“The different URI SHOULD be given by the Location field in the response”), but it was reworded in 7231 and 9110 in a way that assumes the presence of a Location header (“… as indicated by a URI in the Location header field”), while 301, 302, 307 and 308 say “the server SHOULD generate a Location header field”. Well, I reckon See Other with no Location header is fair enough. But URI Too Long was funnier." https://chrismorgan.info/no-query-strings?foo https://chrismorgan.info/no-query-strings?foo
- 1shooner 5mo agoAlso from the 414 page: >Complain to whoever gave you the bad link, and ask them to stop modifying URLs, because it’s bad manners. It's ironic that an error response so blatantly violating the robustness principle is throwing shade about bad manners.
- arjie 5mo agoJust referrer policy of strict origin when cross origin gives host level referer (sic) header in most mainstream browsers unless user has configured otherwise right? That’s usually enough for web authors to know what audience they’re appealing to and privacy-maximizers can turn off that header sending.
- gwern 5mo agoQuery strings break unpredictably, and that alone is enough to ban them by third parties, especially for something as minor as referral tracking. Example: The Browser is a well known link aggregation paid periodical. I subscribe, and every 1 in 10 or 20 links I clicked, it'd just break outright and I'd have to tediously edit the URL to fix it (assuming the website didn't do a silent ninja URL edit and make it impossible for me to remember what URL I opened possibly days or weeks ago in a tab and potentially fix it). This was annoying enough to bother me regularly, but not enough to figure out a workaround. Why? ...Because TB was injecting a '?referrer=The_Browser' or something, and the receiving website server got confused by an invalid query and errored out. 'Wow, how careless of The Browser! Are they really so incompetent as to not even check their URLs before mailing an issue out to paying subscribers?' I wondered the same thing, and I eventually complained to them. It turns out, they did check all their URLs carefully before emailing them out... emphasis on 'before', which meant that they were checking the query-string-free versions, which of course worked fine. (This is a good example of a testing failure due to not testing end-to-end or integration testing: they should have been testing draft emails sent to a testing account, to check for all possible issues like MIME mangling, not just query string shenanigans.) After that they fixed it by making sure they injected the query string before they checked the URLs. (I suggested not injecting it at all, but they said that for business reasons, it was too valuable to show receiving websites exactly how much traffic TB was driving to them on net, because referrers are typically stripped from emails and reshares and just in general - this, BTW, is why the OP suggestion of 'just set a HTTP referrer header!' is naive and limited to very narrow niches where you can be sure that you can, in fact, just set the referrer header.) But this error was affecting them for god knows how long and how many readers and how many clicks, and they didn't know. Because why would they? The most important thing any programmer or web dev should know about users is that "they may never tell you": https://pointersgonewild.com/2019/11/02/they-might-never-tell-you-its-broken/ https://pointersgonewild.com/2019/11/02/they-might-never-tel... (excerpts & more examples: https://gwern.net/ref/chevalier-boisvert-2019 https://gwern.net/ref/chevalier-boisvert-2019 ). No matter how badly broken a feature or service or URL may be, the odds are good that no user will ever tell you that. Laziness, public goods, learned helplessness / low standards, I don't know what it is, but never assume that you are aware of severe breakage (or vice-versa, as a user, never assume the creator is aware of even the most extreme problem or error). Even the biggest businesses.... I was watching a friend the other day try to set up a bank account in Central America, and clicking on one of the few banks' websites to download the forms on their main web page. None of the form PDF download links worked. "That's not a good sign", they said. No, but also not as surprising as you might think - the bank might have no idea that some server config tweak broke their form links. After all, at least while I was watching, my friend didn't tell them about their problem either!
- jedimastert 5mo agoYou know I was actually really curious about this so I went back to the HTML and URL W3C standards and surprisingly they don't actually have any definitions of format other than being percent encoded. One might conflate query strings with "form-urlencoded"[0] query strings, which is one potential interoperability format, but in general a queries string is just any percent encoded string following a "?" in a url[1], and just another property in the "URL" HTML object that can be used in the generation of a response. While additionally there is a URLSearchParams object that is the result of parsing the query string with the form-urlencoded parser, this is simply an interoperability layer for JavaScript. I'm going to be honest, I was pretty geared up to have a contrarian opinion until I looked at the standards but they're actually pretty clear, a 404 could be a proper response to unexpected query string; query string is as much part of the URL API as the path is and I think pretty much everyone can acknowledge that just tacking random stuff onto the path would be ill advised and undefined behavior. [0]: https://url.spec.whatwg.org/#application/x-www-form-urlencoded https://url.spec.whatwg.org/#application/x-www-form-urlencod... [1]: https://url.spec.whatwg.org/#url-class https://url.spec.whatwg.org/#url-class
- nrds 5mo agoWait until you realize that the difference between path and query string is entirely arbitrary and decided by the server. Query strings should never have existed. They are an implementation detail of CGI webservers that leaked all over everything and now smells really bad.
- jolmg 5mo agoIt's arbitrary to a degree like the difference between using an attribute or child element in XML, but it's not entirely arbitrary. If you want to include data in the URL that's not part of the hierarchy of the path, query strings are good for that.
- gpvos 5mo agoQuery strings existed before CGI did, and the way they're defined to be filled in from web forms is quite useful; I wouldn't want to need Javascript to fit that into path format. There's nothing wrong about having things decided by the server; I don't get that part of your argument at all.
- humodz 5mo agoThe tone of this and Chris's post gives me the impression that it's harmful to include these query parameters, but I don't understand how. Could someone elucidate me? I understand it can mangle some URLs and that's good enough reason not do it, but even then it seems like a minor incovenience.
- phoronixrly 5mo agoOh, I have a couple - the users did not agree on being tracked (these query params are tracking information), and the site administrator does not want incoming traffic to be tracked. I know the latter can be hard to understand, but I for example sure as hell do not want to have any info in my logs that can be used to harm my users. On a more personal note, I hate it when I go to copy a link to send via a message, and the tracking code glued onto it is twice as long as original URL... I either have to fiddle around with it to clean it up or leave the person I sent it to to wonder wtf am I on about with a screenful of random characters... So it's violating users' privacy, it's shit UX, and on top of that, nobody asked for it...
- legitster 5mo ago>(these query params are tracking information) Query strings are useful for way more than just tracking. Saving and servicing search queries is a way more common use case. So assuming it's only useful for tracking is very misleading. Query strings are probably the least invasive tracking. They are transparent, obvious, and anonymous. Users are free to strip out and edit query strings if they don't want them. More to the point, I can essentially do the same thing with HTTP routing - create an infinite number of unique URLs for tracking purposes. In that regard calling out query strings specifically for essentially the same thing but more transparently seems like splitting hairs.
- phoronixrly 5mo agoThank you for explaining to me that query parameters can be used for other purposes apart from tracking. The articles in question though, are railing against query parameters being abused for tracking purposes - passing referers (sic) and UTM by adding them to URLs of sites that neither process them, nor want them.
- ChrisMarshallNY 5mo ago> It is a small, decentralised, self-hosted web console that lets visitors to your website explore interesting websites and pages recommended by a community of independent personal website owners. Back in the Stone Age, we called these “Webrings,” but they weren’t as fancy. One of the issues that I faced, while developing an open-source application framework, was that hosting that used FastCGI, would not honor Auth headers, so I was forced to pass the tokens in the query. It sucked, because that makes copy/paste of the Web address a real problem. It would often contain tokens. I guess maybe this has been fixed? In the backends that I control, and aren’t required to make available to any and all, I use headers.
- bch 5mo ago> an open-source application framework, was that hosting that used FastCGI, would not honor Auth headers So you were writing your application as a fcgi-app, and (e.g.) Apache was bungling Auth headers? Can you expand on this? Curious about the technical detail of (I guess) PARAM records not actually giving you what you expect?
- ChrisMarshallNY 5mo agoI don’t remember, exactly. Long time ago (I stepped away from that project many years ago). I just remember the auth headers never showing up in the $_SERVER global (it was a PHP app). This was what I was told was the issue. They made it sound like it was well-known.
- jorams 5mo agoThis is because of a deeply annoying default in Apache, where for "security reasons" the underlying script doesn't get to see auth details that might already be handled by Apache. At some point they added the CGIPassAuth directive[1] but all kinds of other workarounds are floating around on the internet. [1]: https://httpd.apache.org/docs/2.4/en/mod/core.html#cgipassauth https://httpd.apache.org/docs/2.4/en/mod/core.html#cgipassau...
- deleted 5mo ago[deleted]
- legitster 5mo agoQuery strings are awesome. Especially for one-page applications. I build a lot of internal applications, and one of my golden UI rules is that a user should be able to share their URL and other users should be able to see exactly what the sender did. So if you have a dashboard or visualization where the user can add filters or configurations, I have all of their settings saved automatically in the URL. It's visible, it's obvious, it's easy, it's convenient. >There is also a moral question here about whether it is okay to modify a given URL on behalf of the user in order to insert a referral query string into it. I think it isn't. These dogmatic technical screeds are all so weird to me. They usually reveal more about the authors lack of experience or imagination than provide a useful truism.
- keane 5mo agoYes, query strings often enable useful features! But Chris's post, "no unauthorised query strings", is only regarding third parties adding them.
- legitster 5mo agoBut... like... that's a weird hill to die on. > If I wanted to know I’d look at the Referer header; and if it isn’t there, it’s probably for a good reason. You abuse your users by adding that to the link. The reason is that the referrer headers are a usability and privacy nightmare. It's weird for the author to jump to such a conclusion. This referral information is being done purely as a courtesy to the webhost. If we imagined a world in which ChatGPT or Wikipedia launched massive hugs of death on referral links without attributing themselves, that is a much, much worse outcome.
- kyralis 5mo agoThere's a referrer header, if the client wishes to send it. If they don't, the "courtesy to the web host" is done at the expense of the client. This particular web host takes umbrage at other sites taking advantage of their clients that way, which seems reasonable to me.
- 5mo ago
- dang 5mo agoSince the original source hadn't had a discussion on HN yet, I've put that link (https://chrismorgan.info/no-query-strings https://chrismorgan.info/no-query-strings) at the top and moved the response link (https://susam.net/no-query-strings.html https://susam.net/no-query-strings.html) to the toptext. Both are good but it seems fair to give priority to the original.
- ironfront 5mo ago[flagged]
- willthefirst 5mo agoI mean…the site that broke should know what to do with arbitrary query strings. If your site breaks when someone puts in an invalid query string, that’s on you?
- Aardwolf 5mo ago> You could argue that I’m abusing 414 URI Too Long. I respond that it’s funnier this way. Other options I considered were: Another option to consider is "418 I'm a teapot": teapots usually also don't support query strings
- layer8 5mo agoOf course they do. For example you can lower a string from the top to query the fill level. Or you can wrap a string around the pot to query the circumference.
- dredmorbius 5mo agoJust straight "400" ("Bad Request") or "403" ("Forbidden") would also probably be defensible. Odd that there aren't any error response codes specific to URI parameters. Several options which seem like they might be appropriate aren't on close examination: - "406" ("Not Acceptable") which is based on content-negotiation headers. - "409" ("Conflict") which is largely for WebDAV requests. - Others such as 411, 422, and 431 are also for specific conditions which aren't relevant here. - 300 or 500 errors are inappropriate as this isn't a relocation or server-side failure, it's a client-side request problem. Teapot or too long seem best bets.
- mystraline 5mo agoJust fire off a 200 OK with text body of "499 Bad query string" Im not making this up btw. A old NOC I woeked at emitted every error as 200 OK with the body message with the real error. They were a real shitshow.
- deleted 5mo ago[deleted]
- thayne 5mo agoI think either 400 or 404 would be fine. 400 because the request isn't in the expected format, 404 because a resource with that query string doesn't exist.
- notlive 5mo agoReferrer is sometimes nice to know. If your site gets a traffic spike from an email newsletter that traffic won't correctly identify the source in the http headers. No qualms with OP, your site your rules.
- peesem 5mo agoedit: not true https://news.ycombinator.com/item?id=48077990 https://news.ycombinator.com/item?id=48077990 "I don’t like people adding tracking stuff to URLs" and "You abuse your users by adding that to the link" and "no unauthorised query strings" and "At present I don’t use any query strings" but for some reason ?igsh, which i'm pretty sure is an instagram tracking parameter, is allowed. weird
- arexxbifs 5mo agoRunning your own small website is a constant battle against grifters and bad online etiquette. When people hotlink images, I usually make a point of having some personal fun with mod_rewrite.
- dredmorbius 5mo agoThis is genius, kudos Chris. It also makes me wonder what other noxious online behaviours might be addressed through ... creative ... client-side responses similar to this. We've already seen, for years, sites attempting to socially-condition people over the use of ad-blockers and Javascript disablers. No reason why the Other Side can't fight back as well.
- hamdingers 5mo agoWhile I don't take the author's hard stance, I do hate gratuitous query params that result in links that are thousands of characters long. I use this bookmarklet to strip query params before sharing a link: javascript:(()=>navigator.clipboard.writeText(location.origin+location.pathname))();
- kittikitti 5mo agoThis is great! I love one-liners that are readable like this. This made me wonder if there are any extensions that run a script on every page load for a web browser. I'm currently experimenting with userscript managers and plan on including your code as an additional security measure against tracking.
- susam 5mo agoThis corrupts a URL like: https://example.com/?p=20&utm_source=spam to: https://example.com/ when in fact we want the following: https://example.com/?p=20 A possible improvement can be: javascript:(()=>{const u=new URL(location.href);[...u.searchParams.keys()].forEach(k=>{if(k.startsWith('utm_')){u.searchParams.delete(k)}});navigator.clipboard.writeText(u.href)})();
- hamdingers 5mo agoWe have different goals. I do this to strip unnecessary junk off of product page links I'm going to share, and your version now does nothing to (for example) Amazon PDP links. I understand it doesn't work for every kind of link and that's fine. If I did want to support paginated/search pages, I would allowlist only `p` and `q` rather than specifically blocking one type of analytics.
- susam 5mo agoYes, that's fair. Indeed my so called improvement requires continual maintenance to add new parameters to the blocklist, which somewhat defeats what is perhaps the main purpose: convenience. While it is clear that your solution works well for you and would probably work for me too for most types of 'bad' URLs, I am noting down an allowlist based solution for my own satisfaction and future reference: javascript:(()=>{const u=new URL(location.href);[...u.searchParams.keys()].forEach(k=>{if(!['p','q'].includes(k)){u.searchParams.delete(k)}});navigator.clipboard.writeText(u.href)})(); Tested with the following URL: https://duckduckgo.com/?ia=web&origin=funnel_home_website&t=h_&q=hello+world&chip-select=search The bookmarklet copies this cleaned version of the URL: https://duckduckgo.com/?q=hello+world
- shevy-java 5mo ago> It’s my website: I can do what I want with it. > And you can do what you want with yours! That does not make a lot of sense. Yes, you can do what you want with your website, but query-string is a way for users to query for additional information or wants or needs. I use them on my own websites to have more flexibility. For instance: foobar.com/ducks?pdf That will download the website content as a formatted .pdf file. I can give many more examples here. The "query strings are horrible" I can not agree with at all. His websites don't allow for query strings? That's fine. But in no way does this mean query strings are useless. Besides, what does it mean to "ban" it? You simply don't respond to query strings you don't want to handle. We do so via general routing in web-applications these days.
- pessimizer 5mo ago> foobar.com/ducks?pdf This isn't relevant when talking about links to his site. This is relevant when talking about links to your site. > Besides, what does it mean to "ban" it? You simply don't respond to query strings you don't want to handle. It means that you're going to get some sort of 400 error when you follow a link to his site with a query string attached to it. He simply will not respond to query strings that he doesn't want to handle, which is all of them.
- creatonez 5mo ago> "query strings are horrible" That's not at all what the article says. You're responding to a weird strawman that doesn't resemble the article's actual point.
- lloydatkinson 5mo agoThis is really cool. My site is hosted by cloudflare, so I guess I could do the same with a cloudflare worker... maybe?
- jameshart 5mo agoThere’s nothing ruder in hypertext etiquette than giving someone a link to navigate to someone else’s HTTP server, where you have manipulated that URL in some way unsanctioned by the server you are sending them to. You can’t just send arbitrary query string parameters to a server and assume they will just ignore them. Just like you can’t just remove query string parameters and assume the URL will work.
- gojomo 5mo agoIn fact, you usually can just send arbitrary query string parameters to a server - that's why the behavior is so common, and often useful. Most sites don't mind or break, some sites get value from the behavior in ways hard to replicate in other ways – and those sites that don't like such additions can easily ignore them. And a few lines of code will work better than ineffectually appealing to manners, when the freedom of the web's form of hypertext, and protocols, gives the outlink authors full freedom to craft URLs (and thus requests) however they like.
- jameshart 5mo agoCrafting outbound links with your own additions and handing them out to visitors to your site is similar to the practice of writing someone’s phone number on the door of a bathroom cubicle with ‘for a good time call:’ written above it. You’re handing out someone elses’s contact details, but giving the person you hand them to a completely fabricated expectation for how the interaction will go.
- abecode 5mo agoMy use case for this is making separate bookmarks in different folders for a single URL: Example.com/interesting -> bookmark folder one Example.com/interesting?dummy=t -> bookmark folder two
- jameshart 5mo agoUse #fragment identifiers then
- itopaloglu83 5mo agoYouTube is also quite famous with their source identifiers, especially with the short urls, the tracking part is longer then the url I’m trying to share.
- madprops 5mo ago>Right click a youtube video from the results to copy the URL. I would have liked a short URL ready to share with people in chats, but no, I get: https://www.youtube.com/watch?v=IFfLCuHSZ-U&pp=ygUNcmF0Ym95IGdlbml1cw%3D%3D https://www.youtube.com/watch?v=IFfLCuHSZ-U&pp=ygUNcmF0Ym95I... >Want to share an amazon product on a chat to discuss about it. I would have liked a nice short url that I can copy, instead I get a monstrosity, it forces me to manually select only the id portion of it if I want to share it.
- gpvos 5mo agoThis is not the first site to do so. A few years back, scarygoround.com started blocking query strings, although it seems to have stopped doing so now. Back then, Facebook had started to add ?fbclid=... to every outgoing link.
- gojomo 5mo agoTrying to boostrap some taboo against novel unpermissioned URL munging is silly prudishness. Ensuring both sides of a hyperlink agree/consent was a design flaw that limited the uptake of pre-web hypertext systems. The web's laissez-faire approach demonstrated a looser coupling was far better for users, despite all the new failure modes. Of course any site/server has the practical power free to treat inbound requests as rigorously (or harshly) as they want. But by the web's essential nature, it is equally part of the inherent range-of-freedom of outlink authors to craft their URLs (and thus the resulting requests) however they want. URLs are permissionless hyperlanguage, not the intellectual property of entities named therein. Plenty of sites welcome such extra info, and those that don't want it can ignore it easily enough – including by just not caring enough about the undefined behavior/failures to do nothing. Though, when a web publisher has naively deployed a system that's fragile with respect to unexpected query-string values, they should want to upgrade their thinking for robustness, via either conscious strictness or conscious permissiveness. Thereafter, their work will be ready for the real web, not a just some idealized sandbox where scolding unwanted behavior makes sense.
- codingclaws 5mo agoI was just wondering if I should do something like this. I use a couple query string values and I validate them and issue a 40x if the value is invalid. So, I was wondering if I should issue a 40x for an unused query string val.
- huflungdung 5mo ago[dead]
- dspillett 5mo agoMaybe an alternative would be to inconvenience people following such links still, but somewhat less. Instead of responding with an error, give a page that states “The link you followed to get here appears to have had some tracking gubbins added, in case you are a bot following arbitrary links, and/or using random URL additions to look like a more organic visit, please wait while we run a little PoW automaton deterrent before passing you on to the page you are looking for.” then do a little busy work (perhaps a real PoW thingy) before redirecting. Or maybe don't redirect directly, just output the unadorned URL for the user to click (and pass on to others). This won't stop the extra gubbins being added of course, but neither will the error and this inconveniences potential readers less.
- ashley95 5mo agoBut ?fbclid is not banned?
- deleted 5mo ago[deleted]
- Jimmy0252 5mo ago[dead]
- deleted 5mo ago[deleted]
- donohoe 5mo agoA neat and funny idea - but in the end it is hostile to the users who don’t always control what’s added to links.
- sutterd 5mo agoThis url worked fine: https://chrismorgan.info/no-query-strings#:~:text=So%20I%E2%80%99ve%20decided%20to%20try%20a%20blanket%20ban https://chrismorgan.info/no-query-strings#:~:text=So%20I%E2%... but this one was too long: https://chrismorgan.info/no-query-strings?a=1 https://chrismorgan.info/no-query-strings?a=1
- sutterd 5mo agoDoh! The part past the # does not go to the sever, so that wasn't a longer URL. How about: https://chrismorgan.info/%6e%6f-%71%75%65%72%79-%73%74%72%69%6e%67%73 https://chrismorgan.info/%6e%6f-%71%75%65%72%79-%73%74%72%69...
- abanana 5mo agoIndeed, that's not a query string! The #, and following text, is a fragment, is client-side only, and isn't the subject of the blogpost. Neither is percent encoding, which is just another way to send the exact same path from your browser to the server. Note that it has nothing to do with the length of the URL. That's just the error message he's chosen to use, because "4xx stop pissing about with my URLs" doesn't exist in the spec.
- chrismorgan 5mo ago> percent encoding, which is just another way to send the exact same path This is not true for all characters. Some can only be expressed by percent-encoding, and decoding them will either break things completely (e.g. %20) or change the meaning of the URL (e.g. %2F, %3F in paths). Yes, you can encode x as %78 and it should work identically, and you can decode %78 to x and it should work identically—though in both cases, I reckon there’s a strong case for blocking the request as suspicious, and I will probably start doing that soon. But take these examples of improperly decoding: • /foo%2Fbar/baz.html has path «"foo/bar", "baz.html"». • /foo/bar/baz.html has segments «"foo", "bar", "baz.html"». • /foo%3Fbar/baz?quux has path «"foo?bar", "baz"» and query "quux". • /foo?bar/baz?quux has path «"foo"» and query "bar/baz?quux".
- fragmede 5mo agoIt's just a string though. A project that I'll never get to is a custom webserver so that QR codes can use the smaller characterset, so it can link to a URL with parameters without forcing the larger character set.
- wodenokoto 5mo agoSo my understanding is, he is annoyed that other website adds a query string such as "?ref=origin.com" to links pointing to authors website. How does this benefit the other website? How does this hurt the authors website? I am completely confused about the behavior of both side here. I get that when I run an ad-campaing I want google to add a utm-query string, so I can track which campaign users arrived from - but then the origin and the destination are working together. Here the origin just adds stuff for no reason. Why?
- dewey 5mo agoIf you have a popular website and you add that parameter the target easily sees who sends them traffic, that could be the base of sponsorships / affiliate arrangements for example.
- ihateolives 5mo agoBut you see that anyway from access.log or whatever your server supports and dashboard/analytics shows it anyway? What's the benefit of adding origin to query string?
- GeneralMaximus 5mo agoNot always. Some web pages don't send referrers by making all links rel="noreferrer". Mastodon used to do this by default, though now they've changed their stance. Links opened from non-browser apps don't have any referrer information either. E.g. if somebody shares your link on iMessage, WhatsApp, or Telegram. Email clients may also strip out referrers, but I'm not entirely sure about this one. If people read your work via RSS readers, you'll almost certainly not get any referrers. Unless it's a web-based reader like Feedly. My website gets a lot of traffic marked as "Direct / None" by Plausible. I suspect this is traffic from RSS readers or Mastodon, but I can't be sure. A few times I've considered adding a "?ref=RSS" to all URLs served to RSS readers and "?ref=Mastodon" to everything I post on Mastodon. But like the author of this post, I feel uncomfortable tracking my readers like this.
- zzo38computer 5mo agoQuery strings do have uses (such as for searching files and some other kind of dynamic files), but you shouldn't add them to URLs that should not expect them. So, I agree that they would be right to refuse requests with UTM and other stuff added like that. I think 404 probably makes the most sense as the response if a query string is not expected but is present anyways, although 400 might also be suitable.
- lofaszvanitt 5mo agoIMDB recently went haywire they added these ugly qses into every click on their site, bonkers: ?ref_=nm_ov_bio_lk
- llimllib 5mo ago> curl, for example, seems to illegitimately strip a trailing question mark (could be only for the command line, didn’t test library usage). umm what? I don't know what they're actually sending where they think this, but if you think curl is broken you should re-think that maybe you're the one doing something wrong. Here are some examples showing curl not stripping question marks (obviously), I am very curious what this person was actually seeing $ curl -s 'https://httpbingo.org/get?' | jq .url "https://httpbingo.org/get?" $ curl -s 'https://httpbingo.org/get?path' | jq .url "https://httpbingo.org/get?path" $ curl -s 'https://httpbingo.org/get?path,query=bananas' | jq .url "https://httpbingo.org/get?path,query=bananas" $ curl -s 'https://httpbingo.org/get????' | jq .url "https://httpbingo.org/get????" $ curl -sv 'https://httpbingo.org/????' 2>&1 | grep :path * [HTTP/2] [1] [:path: /????]
- chrismorgan 5mo ago$ curl -s 'https://httpbingo.org/get?' | jq .url "https://httpbingo.org/get" This may require further investigation.
- Groxx 5mo agoMight be shell expansion? zsh uses `?` for filename expansion, others might as well: https://zsh.sourceforge.io/Doc/Release/Expansion.html#Filename-Generation https://zsh.sourceforge.io/Doc/Release/Expansion.html#Filena... Though I forget if any shell does stuff like that in quotes. Or printing oddities.
- chrismorgan 5mo agoNo, it’s definitely curl that’s doing it. $ echo 'https://httpbingo.org/get?' https://httpbingo.org/get? $ python >>> import json >>> import subprocess >>> json.loads(subprocess.run(['curl', '-s', 'https://httpbingo.org/get?'], stdout=subprocess.PIPE).stdout)['url'] 'https://httpbingo.org/get' I’m using curl 8.20.0-3, Arch Linux, x86_64. $ curl --version curl 8.20.0 (x86_64-pc-linux-gnu) libcurl/8.20.0 OpenSSL/3.6.2 zlib/1.3.2 brotli/1.2.0 zstd/1.5.7 libidn2/2.3.8 libpsl/0.21.5 libssh2/1.11.1 nghttp2/1.69.0 ngtcp2/1.22.1 nghttp3/1.15.0 mit-krb5/1.21.3 Release-Date: 2026-04-29 Protocols: dict file ftp ftps gopher gophers http https imap imaps ipfs ipns mqtt mqtts pop3 pop3s rtsp scp sftp smtp smtps telnet tftp ws wss Features: alt-svc brotli GSS-API HSTS HTTP2 HTTP3 HTTPS-proxy IDN IPv6 Kerberos Largefile libz PSL SPNEGO SSL threadsafe TLS-SRP UnixSockets zstd
- xp84 5mo agoI love the hilarious output. He even coded in a special case for just a question mark without any params: https://chrismorgan.info/no-query-strings https://chrismorgan.info/no-query-strings? Never have I seen such a sassy web server
- varenc 5mo agogreat spot! I noticed that his server also doesn't accept URLs ending is a single `/`: https://chrismorgan.info/no-query-strings/ https://chrismorgan.info/no-query-strings/ But instead of the banned query strings message, it just returns a very sassy not-a-404 page. Once again, this is violating a common convention, but there's nothing in the HTTP spec that requires treating these URLs the same. Similarly the site also 404s when you add extra slashes like https://chrismorgan.info///no-query-strings https://chrismorgan.info///no-query-strings digression: I love trying "domain.com//" on various sites. Occasionally it'll trigger weird errors like a 502 or 500.
- chrismorgan 5mo agoYou’re misreading a couple of things. There’s no violation of common convention where you’re pointing. Where dealing with static file servers: For URLs that are supposed to include a trailing slash and the server will find that directory and serve its index.html: it’s customary, though not ubiquitous, to redirect from no-slash to slash. (Some, including popular commercial services, serve the index.html file instead of redirecting to add the slash. This is extremely wrong because it changes the meaning of relative URLs.) But the other way round is not common. My URLs don’t include a file extension, and I think that’s influencing your perception into thinking no-query-strings is logically a directory name. But it’s not, it’s logically a file name, just with the .html removed as unnecessary. Take https://susam.net/no-query-strings.html https://susam.net/no-query-strings.html as an example; Susam is more clearly just serving from a file system than I am, and leaves the “.html” file extension in the URL. Do you expect https://susam.net/no-query-strings.html/ https://susam.net/no-query-strings.html/ to work? I hope not. It’s a 404, just as I’d expect, because there is no directory with that name. > not-a-404 page No, that’s a 404, just a plain old boring 404, same as any other. In fact, it’s the same 404 page I’ve been using since 2019, just with dark mode support added. > extra slashes Ah, now for that I had to go out of my way, because Caddy misbehaves out of the box: https://chrismorgan.info/Caddyfile#:~:text=%40has%5Fmultiple%5Fslashes,404 https://chrismorgan.info/Caddyfile#:~:text=%40has%5Fmultiple... > digression: I love trying "domain.com//" on various sites. Closely related is adding the trailing dot of a fully-qualified domain name: https://example.com./ https://example.com./. I didn’t remember to try this on my new site, but it turns out Caddy won’t talk at https://chrismorgan.info./ https://chrismorgan.info./, so that’s probably good.
- patrickdavey 5mo ago"It’s my website: I can do what I want with it. " Right on! It's so liberating having your own wee corner of the internet.
- nullsanity 5mo ago[dead]
- stickfigure 5mo agoThis basically boils down to "reject any incoming links from facebook, pinterest, chatgpt, linkedin, twitter, reddit, youtube, etc". I guess sure? There's a once-famous guy who shows goatse to all referring links from HN. I guess if you get enough traffic that you can pick which sources you want to allow, that's a good problem to have.
- chrismorgan 5mo agoHow much do platforms mangle people’s links? Figured I’d check the ones you mention. (I was actually mildly surprised to be able to find examples in all of them without needing to log in once. I thought LinkedIn and ChatGPT wouldn’t.) Facebook: no. Pinterest: ?utm_source=Pinterest&utm_medium=organic. ChatGPT: ?utm_source=chatgpt.com. (Aside: wow it’s confidently and atrociously wrong if you ask it about me. Ask it just vaguely enough, and it hallucinates someone clearly inspired by me, but who has done a whole lot of stuff that I haven’t. Ask it more precisely about me, and it gets all kinds of details wrong still. I feel further vindicated in hating this stuff. You made me use ChatGPT for the very first time.) LinkedIn: no. Twitter: no. Reddit: no. YouTube: no. > if you get enough traffic that you can pick which sources you want to allow, that's a good problem to have. Nah, I just don’t care about them. It’s my place, I’m doing things on my own terms. Should I discover it to be causing me problems, I’ll burn that bridge when I come to it.
- Sniffnoy 5mo agoFacebook does, I'm not sure why it didn't in your case. It adds an "fbclid" parameter that is quite long. I just tried it to confirm. Edit: Perhaps it only mangles links for logged-in users? That raises the possibility that some of the others may also only affect logged-in users. (Trying with other ones I'm logged in on: Reddit doesn't mangle (obviously), Twitter doesn't mangle.)
- nc0 5mo agoAs far as I know, platforms like YouTube and Twitter prefers using their own "link shorteners" (t.co etc) to track clicks and other metrics.
- himata4113 5mo ago?referrer=123 still works, so I guess it's selective.
- throw310822 5mo agoWhatever floats your boat
- noduerme 5mo agoMost of the sites that still use GET queries around here are the tax collection sites run by local governments, which pass those variables around after you login the way your mom... uh, it's HN, skip the mom joke. I actually get a lot more annoyed by routing parsers that do the same thing Get requests do only by pretending to be a real URL.
- deleted 5mo ago[deleted]
- takethebus 5mo agoI was just thinking about something like this. Instagram and TikTok are major offenders, not everyone wants their personal info blasted everywhere because they copied a link. It would be great to have an iOS shortcut that automatically removes them, it's something I'm going to look into.
- elAhmo 5mo agoSuch a useless blog post and initiative. Author control his website, and if someone lands there by clicking a link, it is not user's fault.
- casey2 5mo agoYouTube couldn't do that
- andersmurphy 5mo agoYup do what works best for you. I do the opposite I don't support path params. Means your router can just be a simple map/dict. Query strings avoid all the hierarchy/taxonomy problems you run into with path params.
- pgt 5mo agoIf query strings are banned, tracking parameters will simply move to the URL, or use `?t=` & `?h=` if those are whitelisted.
- robertjwebb 5mo agoHell yeah
- seabass 5mo agoThis is rad and made me smile
- dmitshur 5mo ago> It’s my website: I can do what I want with it. > > And you can do what you want with yours! This part in particular stood out to me, and I liked it very much.
- deleted 5mo ago[deleted]
- austin-cheney 5mo agoInstead of throwing up an error response wouldn’t he achieve more desirable results from redirecting the address to the same address but without the query string.
- chrismorgan 5mo agoIt’s a form of protest.
- brazukadev 5mo agothis does not achieve the same result. By throwing an error, the person using that url needs to remove the querystring if they want share the link.
- moth-fuzz 5mo agoI dislike the trend among computer users that query strings == tracking data. They're just one of the many ways a browser request can contain data. They are used for all sorts of things in perfectly valid websites, and it's difficult for me as a developer to just just arbitrarily accept that a certain feature HAS to be used for a certain malicious technique, despite it being common. With more limited copy-paste functionality, tracking parameters could just as easily be put in the request body. Or a cookie/session data. Or, if you're nasty, you could get everything you want to know out of a user via fingerprinting with very little up-front data at all. We as users have basically already been collectively pwned, the solution being to use VPNs, anti-tracking scripts/extensions, encryption, and just plain ol 'stop using services that track you'. The preceding are all leagues better for one's privacy than superstitiously chopping up a URL. Honestly, disallowing other websites from from adding their query strings to your URLs I think is awesome, and I think it makes sense that websites should validate their URLs against their internal APIs only. But, given this was only done for query strings, and not other parts of the request, I still feel like this whole thing establishes more of a taboo than an actual security or privacy guideline.
- wotsdat 5mo ago[dead]
- purpleidea 5mo agoI love this. I'd rather see bans when you get sent from one of the google search redirect links. I know they want to track the exits... Maybe mess those up!
- PunchyHamster 5mo agowhy would you do that instead of cutting it out with a redir ? It's essentially telling user "Sorry, you entered my site from wrong side of the internet FUCK YOU"
- romperstomper 5mo agoI noticed that ChatGPT started to add the utm_source parameters for all links it recommends. For example, for GitHub links. Not sure if this critical or not but it could affect users for sure if URLs with any query parameters are blocked.
- miguel-muniz 5mo agoI've created a bookmarklet that just appends a date query string (?date=05112026) to my current tab's URL. My bookmark manager Raindrop will recognize this as a completely different page, so I can easily create a duplicate entries for the same page. I don't think many people mention it, but Raindrop creates archives of every entry, so I end up with sort of my own personal Wayback Machine. The author has successful thwarted me though.