3 ms·
Direct SSH access can be an invaluable tool for debugging production issues. It is great to be able to SSH in and check the tcp dump, view the running processes
by cakeface 10y ago
Direct SSH access can be an invaluable tool for debugging production issues. It is great to be able to SSH in and check the tcp dump, view the running processes in detail, attach gdb and get a memory dump. These types of tasks will never go away. We have great leaps in orchestrating remote application management but when you get down to issues at the bottom of your stack you'll always need to directly access the machine. I like to tell my developers that "the stack goes all the way down". A bug in Linux networking or even a CPU error is still a bug.
- mitchty 10y agoYep, sometimes you have to debug things live with things like perf/gdb echo c > /proc/sysrq_trigger. If not, well congratulations, but if I got told I can't get the above to debug things I'd start question why this "pet" infrastructure lacks basic debugging ability.
- moondev 10y agoSmart health checks and logging should take care of that and remove the instance automatically. You can also spin up a canary machine to "live" debug. I'm referring to distributed clusters not a single machine taking on all the traffic.
- icebraining 10y agoSmart health checks and logging should take care of that and remove the instance automatically. What if the bug is corrupting data or sending incorrect results to the client? Even if you detect and kill the instance, you still have a problem to fix. And even if you can cleanly kill the instance and redirect the request, you can't avoid the latency hit from having to re-process it. You can also spin up a canary machine to "live" debug. You can reproduce the code and data, but how can you reproduce the exact state? You can't log everything that happens in a machine.