3 ms·
It’s not a questions of being able to reverse. It’s a question of being able to diagnose that one of these changes even was the problem and if so which one.
by admax88qqq 2y ago
It’s not a questions of being able to reverse. It’s a question of being able to diagnose that one of these changes even was the problem and if so which one.
- pbhjpbhj 2y agoI focused primarily on guesswho's "in ways I am unaware of". Your issue appears to be true for any system change. Although, risk will of course vary.
- nehal3m 2y agoIf they can be reversed individually you can simply deduce by rolling back changes one by one, no?
- jstanley 2y agoOnly if you already suspect that this tool caused the problem.
- spenczar5 2y agoSuppose you run a fleet of a thousand machines. They all autotune. They are, lets say, serving cached video, or something. You notice that your aggregate error rate been drifting upwards since using bpftune. It turns out, in reality, there is some complex interaction between the tuning and your routers, or your TOR switches, or whatever - there is feedback that causes oscillations in a tuned value, swinging between too high and too low. Can you see how this is not a matter of simple deduction and rollbacks? This scenario is plausible. Autotuning generally has issues with feedback, since the overall system lacks control theoretic structure. And the premise here is that you use this to tune a large number of machines where individual admin is infeasible.
- pbhjpbhj 2y ago>not only can we observe the system and tune appropriately, we can also observe the effect of that tuning and re-tune if necessary. // Does sound like a potential way to implement literal chaos. Surely it's like anything else, you do pre-release testing and balance the benefits for you against the risks?
- Modified3019 2y agoSounds like you have your answer of “don’t use it” then.
- pstuart 2y agoIn that scenario you could run it on a couple servers, compare and contrast, and then apply globally via whatever management tool you use.
- KennyBlanken 2y agoPresumably one would use autotune to find optimized parameters, and then roll those out via change control, either one parameter at a time, or a mix of parameters across the systems. Alternatively: if you have a fleet of thousands of machines you can very easily do a binary search with them to a)establish the problem with the auto-tuner and then b)which of the changes it settled on are causing your problems. I get the impression you've never actually managed a "fleet" of systems, because these techniques would have immediately occurred to you.
- spenczar5 2y agoCertainly when we managed Twitch’s ~10,000 boxes of video servers, neither of the tasks you describe would have been simple. We underinvested in tools, for sure. Even so, I don’t think you can really argue that dynamically changing configs like this are going to make life easier!
- toast0 2y agoWhen you have a thousand machines, you can usually get feedback pretty quick, in my experience. Run the tune on one machine. Looks good? Put it on ten. Looks good? Put it on one hundred. Looks good? Put it on everyone. Find an issue a week later, and want to dig into it? Run 100 machines back on the old tune, and 100 machines with half the difference. See what happens.
- yourapostasy 2y agoRecord changes in git and then git bisect issues, maybe? Without change capture, solid regression testing, or observability, it seems difficult to manage these changes. I’d like to how others are managing these kinds of changes to readily troubleshoot them, without lots of regression testing or observability, if anyone has successes to share.