10 ms·
"Draining and un-draining hosts is hard." I'd stop right there and fix that, because that's a bullshit reason. Cycling hosts in and out of service is easy unle
by hellozomo 3y ago
"Draining and un-draining hosts is hard."
I'd stop right there and fix that, because that's a bullshit reason. Cycling hosts in and out of service is easy unless you're not doing things properly.
The Linux kernel is simply not designed to be live patched and it's a total hack to try to do it, it will never work 100% of the time, always be a source of uncertainty, and always be expensive in terms of engineering work. Disaster will always be looming.
By contrast, fixing their system for taking hosts in and out of service, so that it's extremely robust and reliable would likely pay big dividends in reliability.
My guess would be that this approach is papering over organizational dysfunction. One team can patch all the kernels but one team can't make all the hosts support proper cycling in and out of service. And no one cares to fix it because there's no real incentive to do so. Only cool hacks and new projects are properly rewarded.
- deleted 3y ago[deleted]
- extr 3y ago> My guess would be that this approach is papering over organizational dysfunction. So what? At large enough scale organization problems are harder than technical ones: if you can fix the former with the latter, that's still a win.
- hellozomo 3y agoBut it's not "fixing" the problem, it's papering over the problem by piling tech debt on top of tech debt. Framing this like it's a good thing is my only objection.
- tysonfurytoo 3y agoIt may be easy to cycle hosts in and out, but it can also be time consuming, apparently. In the article it mentions taking 45 days to patch all hosts. The article also points out that this is too long for security updates. Nothing will work 100% of the time. If their patching mechanism is thoroughly tested and battle hardened, I think the risk would be acceptable. Once you do the initial kpatch security upgrade, you could even schedule the machine for serivce so that it's not relying on that, limiting your exposure to bugs.
- deleted 3y ago[deleted]
- WookieRushing 3y agoWhy do you think its easy? There's a lot of systems where you can easily take down some hosts, but taking down more than N% at a time causes issues. If your fleet is large enough then you are limited by the largest set of hosts where you can only take N% down at a time. Now you could say keep the sets of hosts small or N% large. But that can cause other issues as you typically lose efficiency or zonal outage protection. A solution to this could be VM live migration or something similar. This breaks down for storage systems where you can't just migrate those disks virtually since they're physical disks or places that don't use VMs.
- natbennett 3y agoThe work that makes cycling hosts in and out easy is itself hard. I agree that it’s the right thing to do but it’s hard.
- fragmede 3y agoIt's not a bullshit reason. You can put all the lipstick you want on the pig and put software all around it to make it all sorts of easy, but at the end of the day, having to reboot is a stop-the-(machine's)-world situation. Not having to do that is just better. Even if it doesn't work 100% of the time, that's still better than having to reboot the whole fleet. 45 days to reboot the whole fleet! Throwing FUD and saying disaster is looming because its scary computer magic (out of MIT) was a scare tactic RedHat used to throw around about Oracle/Ksplice until they developed their own (Kpatch), then suddenly their sales team had to backtrack and say actually hot patching is good and can be trusted. I'm not saying it's not risky or dangerous, it's operating in kernel space, but that's why they pay really smart people to be careful when doing it, and not digital equivalent of a plumber who can't do more than glue libraries together. A better understanding of the underlying technology so it's less magic might assuage your fear of it, but thinking Facebook is so dysfunctional that they haven't already made it easier to reboot is to misunderstand the problem at hand.
- deleted 3y ago[deleted]
- mypalmike 3y ago> fixing their system for taking hosts in and out of service, so that it's extremely robust and reliable would likely pay big dividends in reliability. Facebook hosts can be robustly cycled. Of course. They've been doing this stuff for years. They've figured it out. That's not the issue. Scaling up brings about new problems. This article specifically mentions the 45 day rolling restart issue. That's not an issue when you have 1000s of servers. It's one that shows up a couple of orders of magnitude later. So you either solve the problem with a hack like kernel patching or you work to reduce restart times (drain + shutdown + OS restart + process initialization across every service). Get those restart times down 50% (good luck accomplishing that) and congrats, you're down to maybe a 25 day rolling restart, which is still quite a problem.
- hellozomo 3y agoWhat exactly do you think requires cycling millions of hosts to take 45 days? That's a couple hundred hosts every five minutes across a large number of datacenters? I wouldn't expect it to be halved by optimization. I'd expect it to be an order of magnitude faster and take more like 4.5 days. I wouldn't be surprised (but I would be impressed) if they tried and got it down to a full cycle requiring one working day. That's around 1% of hosts cycling every five minutes.
- xyzzy_plugh 3y agoIt's certainly weird. I've worked at ~million host scale where uptime never exceeded a week by default (with the odd carve out for problem child software supplied by vendors) I'm betting it's moreso that teaching developers to write software that tolerates draining properly (or is even able to communicate draining) is too difficult for them so they work around it.
- sangnoir 3y ago> It's certainly weird. I've worked at ~million host scale where uptime never exceeded a week by default How many individual teams had software running on your hosts? How many those hosts were stateful, and were fragmented across hundreds or thousands of service groups that had their own fault tolerances and unknown (to infra team) warm-up times. Adding complexity (rolling reboots) to already complex systems is almost never a good idea - at some point, there will be an issue caused by hosts rebooted in the wrong order, or too many hosts of a certain type 2-dependency-levels down being simultaneously offline
- raincom 3y agoRedhat provides kpatches for 6 months, that's all. If you are running a year old kernel, no kpatches are provided for that kernel. Definitely, one needs to recycle hosts every six months.
- lokar 3y agoAt google we did pretty much the same thing. Aimed to be able roll a kernel in 30 days, but various edge cases always made it drag out at the end unless you really spend a lot of human time on it. So use kaplice for really critical stuff (where the patch was easy, not always the case). A reasonable compromise in the real world.
- ungamedplayer 3y agoAt redhat, we just maintain kpatch and hire kernel engineering to do the same. Turn around is about a week for 40 variants of kernel.
- jb_gericke 3y agoYeah, especially with containerisation and orchestration / Kubernetes, I get that perhaps not everything is viable to containerise, but in 2023 this feels archaic and like a lot of (potentially unnecessary) engineering work.