4 ms·
So, with No SSH - how do you debug that one-off problem that is only on machine abc12? I'm not talking about mutating, I'm talking about attaching gdb to the pr
by bluecmd 11y ago
So, with No SSH - how do you debug that one-off problem that is only on machine abc12? I'm not talking about mutating, I'm talking about attaching gdb to the process while it's handling requests. I'm talking about collecting CPU profiling information from production. Stuff like that.
- mhw 11y agoImmutable Infrastructure is the modern day "have you tried turning it off and back on again?"
- autotune 11y agoYou could probably have a central logging service like graylog or logstash that automatically collects that information and just sort through the logs that way.
- AznHisoka 11y agoWhat if that 1 server ran out of file handles and can no longer log to anywhere?
- deleted 11y ago[deleted]
- xxpor 11y agoTake the server out back and shoot it.
- autotune 11y agoLaunch a copy of that server with a new image (which I would hope is under a load balancer), update the sysctl.conf file template with a larger fs.file-max limit in whatever config management system you use for that instance, resume logging under freshly cloned instance.
- dsp1234 11y agoHow would you know to do that at all? The proposed issue is that the file handles are exhausted. Thus logging doesn't actually work, and thus you don't know anything about the problem except 'it's not working'. Your solution presupposed that you have figured out the issue, but the question is how do you find out about the problem at all. Without SSH/debugger/whatever, what is the process to inspect the system to find the problem?
- autotune 11y agoGreat question. I'd probably look into something like Sysdig Cloud that can collect that information through its own agent, or one of the tools mentioned by seanp2k2. I do also agree with his point on edge cases though, it should be the goal for the majority of your infrastructure while still allowing for edge cases.
- devonkim 11y agoAfter dealing with different enterprise systems in AWS for a few years now, I'd say that this is what AWS ELB feature to remove a system out of an ASG as part of its lifecycle exists for. If you aren't using similar horizontal scaling strategies, you have a disaster waiting to happen basically and basically deserve an outage you can't diagnose / RCA well. If you can't ssh in afterwards that means your provisioning for automation / instrumentation and you should stay be able to isolate attention to that part of an instance's lifecycle.
- api 11y agoWhat if the problem is not diagnosable from logs? In a perfect world production and development would be identical, all production issues would be easy to reproduce in development, and logging and error handling would be sufficient to diagnose any problem. But what do you do when you have a problem that doesn't happen in a perfect world -- that you can't diagnose normally, doesn't occur in dev, shows nothing abnormal in logs, etc.? At some point you need the ability to inspect the actual thing. It's like diagnosing a problem with a bridge by doing experiments on a scale model replica. Doesn't always work.
- autotune 11y agoIn that case I'd agree with you, and you should still allow for the possibly of logging in via ssh, while aiming to reduce that need and focus on automation as much as possible, or as much as it's appropriate and possible for your environment, rather.
- seanp2k2 11y agohttps://speakerdeck.com/niteshkant/distributed-tracing-at-netflix https://speakerdeck.com/niteshkant/distributed-tracing-at-ne... http://techblog.netflix.com/2012/06/scalable-logging-and-tracking.html http://techblog.netflix.com/2012/06/scalable-logging-and-tra... https://github.com/Netflix/vector https://github.com/Netflix/vector Not terribly easy, but possible. I agree that "never have to SSH" is a bad target because there will always be edge cases, but for the majority of things it's possible to troubleshoot at the cluster level given immutable infrastructure; all the systems will be doing the exact same thing given an input. Reality shows that this is a bit idealistic, but if you start finding that e.g. A certain set of requests cause weird behavior, you can look at what makes those requests different, then set up some routing logic to send them to a debug farm (using something like https://github.com/Netflix/zuul https://github.com/Netflix/zuul ).