4 ms·
The idea that an "observability stack" is going to replace shell access on a server does not resonate with me at all. The metrics I monitor with prometheus and
by crawshaw 9mo ago
The idea that an "observability stack" is going to replace shell access on a server does not resonate with me at all. The metrics I monitor with prometheus and grafana are useful, vital even, but they are always fighting the last war. What I need are tools for when the unknown happens.
The tool that manages all my tools is the shell. It is where I attach a debugger, it is where I install iotop and use it for the first time. It is where I cat out mysterious /proc and /sys values to discover exotic things about cgroups I only learned about 5 minutes prior in obscure system documentation. Take it away and you are left with a server that is resilient against things you have seen before but lacks the tools to deal with the future.
- gear54rus 9mo agoAgreed, this sounds like some complicated ass-backwards way to do what k8s already does. If it's too big for you, just use k3s or k0s and you will still benefit from the absolutely massive ecosystem. But instead we go with multiple moving parts all configured independently? CoreOS, Terraform and a dependence on Vultr thing. Lol. Never in a million years I would think it's a good idea to disable SSH access. Like why? Keys and non-standard port already bring China login attempts to like 0 a year.
- ValdikSS 9mo ago>What I need are tools for when the unknown happens. There are tools which show what happens per process/thread and inside the kernel. Profiling and tracing. Check Yandex's Perforator, Google Perfetto. Netflix also has one, forgot the name.
- reactordev 9mo agoOr… you build a container, that runs exactly what you specify. You print your logs, traces, metrics home so you can capture those stack traces and error messages so you can fix it and make another container to deploy. You’ll never attach a debugger in production. Not going to happen. Shell into what? Your container died when it errored out and was restarted as a fresh state. Any “Sherlock Holmes” work would be met with a clean room. We have 10,000 nodes in the cluster - which one are you going to ssh into to find your container to attach a shell to it to somehow attach a debugger?
- toast0 9mo ago> We have 10,000 nodes in the cluster - which one are you going to ssh into to find your container to attach a shell to it to somehow attach a debugger? You would connect to any of the nodes having the problem. I've worked both ways; IMHO, it's a lot faster to get to understanding in systems where you can inspect and change the system as it runs than in systems where you have to iterate through adding logs and trying to reproduce somewhere else where you can use interactive tools. My work environment changed from an Erlang system where you can inspect and change almost everything at runtime to a Rust system in containers where I can't change anything and can hardly inspect the system. It's so much harder.
- IgorPartola 9mo agoSay you are debugging a memory leak in your own code that only shows up in production. How do you propose to do that without direct access to a production container that is exhibiting the problem, especially if you want to start doing things like strace?
- joshuamorton 9mo agoI will say that, with very few exceptions, this is how a lot of $BigCo manage everyday. When I run into an issue like this, I will do a few things: - Rollback/investigate the changelog between the current and prior version to see which code paths are relevant - Use our observability infra that is equivalent to `perf`, but samples ~everything, all the time, again to see which codepaths are relevant - Potentially try to push additional logging or instrumentation - Try to better repro in a non-prod/test env where I can do more aggressive forms of investigation (debugger, sanitizer, etc.) but where I'm not running on production data I certainly can't strace or run raw CLI commands on a host in production.
- reactordev 9mo agoCombined with stack traces of the events, this is the way. If you have a memory leak, wrap the suspect code in more instrumentation. Write unit tests that exercise that suspect code. Load test that suspect code. Fix that suspect code. I’ll also add that while I build clusters and throw away the ssh keys, there are still ways to gain access to a specific container to view the raw logs and execute commands but like all container environments, it’s ephemeral. There’s spice access.
- ValdikSS 9mo ago>It is where I attach a debugger, it is where I install iotop and use it for the first time. It is where I cat out mysterious /proc and /sys values to discover exotic things about cgroups I only learned about 5 minutes prior in obscure system documentation. It is, SSH is indeed the tool for that, but that's because until recently we did not have better tools and interfaces. Once you try newer tools, you don't want to go back. Here's the example of my fairly recent debug session: - Network is really slow on the home server, no idea why - Try to just reboot it, no changes - Run kernel perf, check the flame graph - Kernel spends A LOT of time in nf_* (netfilter functions, iptables) - Check iptables rules - sshguard has banned 13000 IP addresses in its table - Each network packet travels through all the rules - Fix: clean the rules/skip the table for established connections/add timeouts You don't need debugging facilities for many issues. You need observability and tracing. Instead of debugging the issue for tens of minutes at least, I just used observability tool which showed me the path in 2 minutes.
- crawshaw 9mo agoHow did you use tracing to check the current state of a machine’s iptables rules?
- ValdikSS 9mo agoIn this case I used `perf` utility, but only because the server does not have a proper observability tool. Take a look at this Netflix presentation, especially on the screenshots of their web interface tool: https://archives.kernel-recipes.org/wp-content/uploads/2025/01/kernelrecipesperfevents-170929090404-171007142608.pdf https://archives.kernel-recipes.org/wp-content/uploads/2025/...
- crawshaw 9mo agoThat is a command line tool run over ssh. If you have invented a new way to run command line tools, that’s great (and very possible, writing a service that can fork+exec and map stdio), but it is the equivalent to using ssh. You cannot run commands using traces.
- deleted 9mo ago[deleted]
- jeffbee 9mo agoI guess the question is why your observability stack isn't exposing proc and sys for you.
- crawshaw 9mo agoMine (prometheus) doesn’t because there are a lot of high-dimensional values to track in /proc and /sys that would blow out storage on a time-series database. Even if they did though, they could not let me actively inject changes to a cgroup. What do you suggest I try that does?
- jeffbee 9mo agoExperience from another company where I (and you) worked suggests that having the endpoints to expose the system metrics, without actually collecting and storing them, is the way to go.
- crawshaw 9mo agoYears of debugging in that company’s restricted environments solidified my desire for shell access to production environments. I was there a month before I was hunting for breadcrumbs in a BINARY_INFO log that I had five minutes to grab before it was deleted.
- jeffbee 9mo agoWell that's funny you mentioned it because one of my projects was a service that let users temporarily install binary info logs collectors triggered by predicates, remotely, which at least I thought was a better model than ssh into the host or, for the advanced caveman, pdsh into many hosts. I don't really see a reason why I can't do that for gRPC, either ... But, anyway, remote command and control of observability really is a thing in the industry, not just at one company.
- cyberax 9mo agoBecause you're holding it wrong! The dashboards are something that looks cool, but they usually are not really helpful for debugging. What you're looking for is per-request tracing and logging, so you can grab a request ID and trace it (get log messages associated with it) through multiple levels of the stack. Even maybe across different services. Debuggers are great, but they are not a good option for production traffic.
- raggi 9mo agoYep. Observability stacks are a similar blind alley to containers: They solve a handful of defined problems and immediately fall down on their own KPI's around events handled/prevented in-place, efficiency, easier to use than what came before.
- cryptonector 9mo agoThe problem lies in surveillance and others understanding what you did. Say your security department records every shell interaction with prod services: how does one then review and understand what happened? This is a fairly tricky problem. Perhaps through it at an LLM, but it'd have to be well trained to look for malicious actions.