5 ms·
Even though a GCE instance can be made to boot pretty quicky it's a bit sad that it takes around 40-50 seconds before any network connectivity can be establishe
by andoma 9y ago
Even though a GCE instance can be made to boot pretty quicky it's a bit sad that it takes around 40-50 seconds before any network connectivity can be established.
I asked a question about it here a while ago https://serverfault.com/questions/845298/no-network-connectivity-until-one-minute-after-boot https://serverfault.com/questions/845298/no-network-connecti...
Checked again just now, approx. 40 seconds between pressing START in the console UI and until I can ping 8.8.8.8, using serial console on Debian Stretch.
If anyone have a clue how to improve this I'd be most happy to hear about it. I'm using this for CI builds.
- boulos 9y agoWe're working on it. It's a major initiative for us, because as you see we get the damn thing "booting" in a handful of seconds and then reachability to/from the internet is the long pole. Fwiw, we at least got to/from *.googleapis.com way down, so if you need to say fetch something from GCS, that should be a bit faster. Disclosure: I work on Google Cloud.
- andoma 9y agoThanks, Very happy to hear this :-)
- _asummers 9y agoCurious (if you can go into it) what the current technical hurdles are in getting the connectivity times down, and/or what you did to get the internal connectivity going quickly.
- boulos 9y agoAt a high level, it's the result of having global, flat Networks and not wanting to declare the network "up" until you've "programmed" all the routes. So if you have 1000 VMs distributed globally, you get to make sure that they're not "connected" until your new VM in asia-east1-a can talk to all other VMs in your Network (and vice versa). With the to/from API path this routing is much simpler since you don't get the N^2 behavior. Again, Disclosure: I work on Google Cloud (but I've never contributed engineering-wise to Cloud networking).
- _asummers 9y agoWould that connectivity happen regardless of firewall rules and such (thinking network tags specifically)? Like if I have 1000 VMs as in your example, and I spin up a new box but have a tule that says only one of those 1000 can talk to it, would the route programming still go through and assert that connectivity to all 1000? Or would it be smart enough to realize that only one of those needed to connect and finish more quickly? If the latter, what happens if I then remove that firewall rule immediately after boot?
- jsolson 9y agoToday's control plane is fairly smart about which routes are required for a given config change event (there's a lot one could speculate about here, especially around the word "required") -- fwiw, I also don't work on the control plane -- I have spent a lot of time on the GCE hypervisor network dataplane. So a fair amount of hand waving follows -- just assume details are missing because I don't know what they are :) We aim for global convergence of network state as sort of an ongoing goal, but it's a distributed system with failure domain isolation, so that goal is necessarily flexible. There are different rates of convergence, and Solomon is certainly right that routes to first party services are some of the easiest to converge. Internet connectivity is to a certain extent the hardest thing to converge, as in our premium tier (which up until recently was our only tier) we aim to keep data on Google's network for as much of its journey as possible. At the extreme, this means a lot of edge nodes learning that a given external IP belongs to your VM. Part of it is also just the mundane business of reconciling what a given configuration event means and propagating that to interested parties. With respect to your firewall example, it's really just another config change event with some set of implications for routes that are added or removed. Anyway, that's my hand wavy explanation. Hopefully it's helpful!
- Lukas_Skywalker 9y agoOT: boulos/jsolson, thanks for being so open about the platform. Getting information like this has become rare, but it makes a very interesting read.
- ikiris 9y agoWow really? Why does the stack take that long to establish end to end? (Feel free to answer me internally :) )
- IanCal 9y ago> Fwiw, we at least got to/from *.googleapis.com way down, so if you need to say fetch something from GCS, that should be a bit faster. Not used to GCE, what sort of timeframe is this expected in? Kinda wondering what the expected time would be for a process that's something like: Run container Download something from GCS null op Upload to GCS Terminate I know there's a lot of variables there, but does anyone have a quick idea of what this should be?