6 ms·
I've been here with this and having logs also emitted by application/web servers is critical too. If you only have it at the gateway, and it emits a 503 or 504,
by rdoherty 3y ago
I've been here with this and having logs also emitted by application/web servers is critical too. If you only have it at the gateway, and it emits a 503 or 504, did the request make it to the web server? Maybe? You have no signal. Seeing a timeout at the gateway/load-balancer and the web server log that it serviced the request successfully but took 45s tells you very critical information. If you didn't have the web log, you have a missing signal that tells you critical information.
Despite having to managed TBs of logs per day and sift through them at times, I'd rather have too much logs. However, we did not alert on any of them. All alerting was symptoms based via SLOs (error rate, latency). Logs were only used for debugging.
- teaearlgraycold 3y ago> If you only have it at the gateway, and it emits a 503 or 504, did the request make it to the web server? Maybe? I was taught a perfect solution to this. Only ever return 200-300 from your web server (not applicable for most public APIs, though). My web server is not a gateway, not a resource access mechanism as envisioned by people in the 90s. It’s always an RPC server. REST is RPC. JSON is RPC. I let my outsourced CDN do HTTP code shenanigans, but any apps I develop are 200 all the time. Want to know if it failed? Easy! {“success”: false, …}
- nesarkvechnep 3y agoThat’s why we can’t have nice things.
- teaearlgraycold 3y agoI and many others have done it for years. HTTP codes were not designed for the web in 2023. They were for the 90s.
- dragonwriter 3y agoNothing has changed which would invalidate the basic model, though. “It’s old” isn’t a reason to throw out a working design.
- teaearlgraycold 3y agoNo, the reason is "It was built for a time when people expected the web to mostly be a series of interconnected static documents"
- dragonwriter 3y agoThat still doesn’t explain any actual problem, it explains what you think is the cause behind the unspecified problem, jumping right over the key question.
- teaearlgraycold 3y agoBasically, HTTP is the envelope. Ethernet > IP > TCP > HTTP > My application HTTP is a hyper text transfer protocol, designed as such. My application is a RPC front-end for a series of data stores, AI models, etc. Using HTTP codes for your application's protocol is a mismatch. There are some fairly general-purpose HTTP error codes that can meet many application's use cases, but the overlap is not perfect. You can get greater granularity by ignoring it entirely. Going HTTP 200-only means any non-200 is now definitively from 3rd party middleware, a router, proxy, etc. The boolean check of "did I hit my application code?" is easy and useful.
- dragonwriter 3y agoIts easy and useful if you just return details in response bodies, which you can do with errors as well as 200s, and that lets everything else (including client code) that deals with your responses (which doesn’t generally care if it hit your application code as likely as it cares about the things HTTP error codes communicate) handle errors without having to code for your bespoke method of error reporting.
- gwbas1c 3y ago2xx: Request succeeded 4xx: You screwed up 5xx: I screwed up Easy! More important, this is generally expected when calling any kind of API over HTTP. Returning a 200 in an error situation makes it very difficult to diagnose errors. For example, the F12 debugging pane in my browser color-codes errors responses. Fiddler does the same thing too. More important: Languages / libraries have built-in mechanisms for handling these errors, which make error situations easier to handle (or not handle.) For example, some C# libraries will always throw an exception on a 4xx or 5xx response. If your oddball API returns a 200 in an error situation, it breaks normal error handling patterns. (And could result in making situations where "success": false much harder to track down.) And finally: Because error situations are unusual, they often aren't encountered in the normal day-to-day debugging of your APIs consumers. As a result, your oddball API's failure situation is probably untested in whoever is calling your API. When your API fails (because it will at some point,) there's a high chance your consumer won't check the success flag and parse the response as if it's successful. (Again, because you're not following normal HTTP semantics.) This will trigger unpredictable failures in your consumers.
- teaearlgraycold 3y ago> Languages / libraries have built-in mechanisms for handling these errors And it's always easy for me to override this. I've been doing it this way for years - all of the startups I've worked at do it this way - and it's extremely successful. > whoever is calling your API Just internal use. For external use I agree meeting expectations is important. Users will expect a 400/500 for certain issues.
- gwbas1c 3y agoRemember, you're breaking debugging tools like Fiddler when you do this... ...And, whoever inherits your code probably won't like working with it. Doing weirdo things like that can hurt your reputation in a team.
- lelanthran 3y ago> Remember, you're breaking debugging tools like Fiddler when you do this... > ...And, whoever inherits your code probably won't like working with it. Doing weirdo things like that can hurt your reputation in a team. Well, it explains why web APIs are such a hot mess: the people creating the tools don't know the difference between transport errors and application errors, the developers creating the applications don't know the difference, and the developers calling into the application don't know the difference. It's a pretty important difference, and it all other protocol stacks care is taken not to side-step any layer and directly fiddle with the transport layers from the application layers. Right now, a client getting a 5xx response can't tell if the application had an error or if the proxy is misconfigured, because the application developers are sending proxy errors/server errors (5xx) and trampling all over the namespace of the transport medium. The system appears to be well structured, but the conventions were all set by developers who were all noobs. I don't think, in 25 years of development, I ever saw a sockets program (neither client nor server) where the application detected an error and set error-bits in the IP datagram, and yet in web-development, I see this all the time.