3 ms·
What exactly do you think requires cycling millions of hosts to take 45 days? That's a couple hundred hosts every five minutes across a large number of datacent
by hellozomo 3y ago
What exactly do you think requires cycling millions of hosts to take 45 days? That's a couple hundred hosts every five minutes across a large number of datacenters?
I wouldn't expect it to be halved by optimization. I'd expect it to be an order of magnitude faster and take more like 4.5 days.
I wouldn't be surprised (but I would be impressed) if they tried and got it down to a full cycle requiring one working day. That's around 1% of hosts cycling every five minutes.
- xyzzy_plugh 3y agoIt's certainly weird. I've worked at ~million host scale where uptime never exceeded a week by default (with the odd carve out for problem child software supplied by vendors) I'm betting it's moreso that teaching developers to write software that tolerates draining properly (or is even able to communicate draining) is too difficult for them so they work around it.
- sangnoir 3y ago> It's certainly weird. I've worked at ~million host scale where uptime never exceeded a week by default How many individual teams had software running on your hosts? How many those hosts were stateful, and were fragmented across hundreds or thousands of service groups that had their own fault tolerances and unknown (to infra team) warm-up times. Adding complexity (rolling reboots) to already complex systems is almost never a good idea - at some point, there will be an issue caused by hosts rebooted in the wrong order, or too many hosts of a certain type 2-dependency-levels down being simultaneously offline
- xyzzy_plugh 3y agoI appreciate your attempt to invalidate my experience but your points are irrelevant. > hosts rebooted in the wrong order Order doesn't matter. Host groups set a threshold for unavailability. Hosts are not rolled unless availability targets are maintainable. Usually this just means the oldest host at any time will get rolled. Facebook obviously uses hot patching kernel updates to work around a social issue. Instead if you are functionally able to prescribe a set of behaviors that teams must comply to, you can easily do things like rebooting the fleet monthly without impacting availability regardless of the statefulness or fault tolerances. If I shoot a random host in your pool and it matters to you then you haven't achieved fault tolerance. I'm obviously not proposing shooting an unfair number of hosts to you.
- sangnoir 3y ago> I appreciate your attempt to invalidate my experience but your points are irrelevant That was not my intention - I genuinely would have appreciated answers to my questions as it be useful to compare the complexity of your setup versus Facebook. As an extreme case: million homogenous, stateless hosts are far less complex to manage compared a million heterogeneous, stateful ones, and very little translates from the former to the latter - in my experience. > Facebook obviously uses hot patching kernel updates to work around a social issue. Which I think is reasonable when you have tens of thousands of SDEs. > I'm obviously not proposing shooting an unfair number of hosts to you. I agree with you, but I'll go on to say "not shooting an unfair number of hosts" is a hard problem to solve at scale, unless you're willing to make it simple and make humans deal with it by continually draining/undraining services which costs a lot of money without increasing the top line, likely far more money than it cost to get a handful of engineers to write kernel splicing. So beyond it being possibly a social issue, it may be a cost/host utilization issue as well
- xyzzy_plugh 3y agoThese are fair points but I'd add that continually draining amortizes to $0 as the fleet grows. Even if you can splice there are benefits to limiting uptime, with maintenance reaping the majority.
- mypalmike 3y agoThe amount of time needed to restart a fleet (without sacrificing availability) is correlated with excess server capacity. Excess server capacity is not free.
- raverbashing 3y ago1 Million hosts over 45 days = 15 min per host That's a very realistic/optimistic number (especially as you do want to wait for all services to be running and marked as healthy) "Oh but you can batch this" sure, but you don't want too much of a big batch that will make your service slow or want to risk shooting yourself in the foot - like rebooting your whole control plane then figuring out it doesn't work like that (The 45 days is probably an estimate as well, I'm not sure they actually do that server by server)
- exitheone 3y agoThat's still ridiculously slow. I'd expect them to have hundreds of Microservices. Each one of those should be able to handle a random restart at any point in time so they should absolutely be able to restart 100s of servers concurrently without major disruptions. Hell on Facebook scale a whole-Datacenter going down should not cause service disruptions.
- Closi 3y agoThis does assume that nothing is getting broken along the way. Taking 45 days is probably more about caution and resolving issues systematically rather than pushing a big button and hoping you don’t cause issues. I’d expect them to have thousands of microservices - and you only have to find a way to break one to cause big issues.
- exitheone 3y agoRegular random crashes should be exercised regardless at Facebook scale. Not being resilient to that would be very unprofessional.
- BlackFingolfin 3y agoThat's not 15 minutes per host; that's 15 hosts per minute.