3 ms·
How Kubernetes Probes Work
- sidcool 2mo agoThis does not state anything new, but explains it so much well than the kubernetes documentation.
- srichard16 2mo agoSometimes the k8s docs remind me of google's documentation
- leetrout 2mo agomany times i have said a company could be built to just make better google documentation.
- dev_cprice 2mo agoSam has a real way with words when it comes to educational content. Though he did find a legit Kubernetes bug while writing the post, so technically there was at least one new thing :)
- stackskipton 2mo agoSRE here, Strong disagree with do not fail readiness and liveness checks on upstream dependencies failing. There are several reason to do so and unless you have extreme start up time, what's the problem with restarting? Maybe DNS has changed on you but you are stuck with bad local cache because you poorly respect TTLs (Looking at you Java), reseting the process will clear that cache away. Maybe TCP connections are in stuck weird state, resetting the process generally helps with that. Maybe someone gave you bad ENV VARs and you cannot connect to database, by refusing to progress the rollout, no outage generated. So yea, if you are not ready to do work including critical upstream dependencies, don't lie to system and say you are.
- arccy 2mo agoThundering herd / cascading outages. You take out a large enough portion of your fleet, and the remaining load overloads your remaining nodes one by one as they restart, so you can never have enough healthy nodes.
- erulabs 2mo agoSRE team debates correctness versus availability for the 540th time this year You're both correct, of course!
- atmosx 2mo agoWhat this guy said :point_up: My personal take-away is this: whatever you choose, make sure it's consistent across services (not serviceA behaves like X and serviceB like Y) and make sure eng teams know _how_ these are configured and what can go wrong. They'll figure out the rest.
- solatic 2mo agoYou'd be surprised how many engineering leaders don't understand the CAP theorem and will fail engineers on interviews for picking the one they don't agree with instead of communicating their expectations clearly (dodged a bullet on that one ...)
- jaggederest 2mo agoThat's a problem for circuitbreakers on these kinds of actions, not lying on health checks. Something like healthcheck fails -> restart -> healthcheck fails -> restart -> healthcheck fails -> circuit breaker trip, alarm raised, give up until manual intervention or X minutes have passed
- deathanatos 2mo agoThat circuitbreaker exists, by default. It is "CrashloopBackoff", here, and TFA covers it. (& it's an "until X minutes have passed" kind, by default.)
- stroebs 2mo agoI need to know how to animate things like this for internal documentation.
- claytonjy 2mo agoThe author re-implemented a bunch of k8s logic in a typescript library just for these animations: https://github.com/ngrok/webernetes https://github.com/ngrok/webernetes
- smartbit 2mo agoI ported Kubernetes to the browser https://ngrok.com/blog/i-ported-kubernetes-to-the-browser https://ngrok.com/blog/i-ported-kubernetes-to-the-browser 3 comments https://news.ycombinator.com/item?id=48734656 https://news.ycombinator.com/item?id=48734656