4 ms·
An interesting side-effect of having awesome tooling. When you have a service that is taking too much memory and can see directly on a Grafana dashboard that th
by cfors 7y ago
An interesting side-effect of having awesome tooling. When you have a service that is taking too much memory and can see directly on a Grafana dashboard that the reason your service keeps getting OOM'd because of too many Goroutines you can avoid even having to go into a box to debug.
When I used to be on call, I remember very distinctly an incident where there was a ton of restarts on containers for a set of nodes (This was on K8s, so not immediately obvious). The operations team mentioned to me that in the logs (K8s logs here) there was a lot of errors being generated along the lines of "Discovery failing, restarting".
I asked if we had a tcpdump if we could see what was happening, and the answer I got was "There is no tcpdump in our container." Meaning that we didn't have the binary in the actual container image.
For anyone with any knowledge of Kubernetes, you know that you can easily just SSH to the machine running the pod, get the current port it's being ran on and then run a tcpdump on there. However, the fact that all of this couldn't be tied together to come up with a tcpdump to understand that caused this issue to persist for an extra 3 hours when the issue was a flapping NIC that was misconfigured.
Without understanding how your systems actually work under the hood, you will be running them with your hands tied. It's not good enough to understand one abstraction, you need to be prepared at anytime to peel back the layers of abstraction and get your hands dirty in a different layer.