3 ms·
My guess is that whatever clever network optimizations that Google has are probably interfering with their traffic. By building their own network stack, they
by devsda 3y ago
My guess is that whatever clever network optimizations that Google has are probably interfering with their traffic.
By building their own network stack, they are skipping them and also wireguard might be better equipped to dealt with occasional faults as it built on udp which is inherently unreliable.
- Bluecobra 3y agoI have a direct cross connection to Google in a colocation facility (aka Dedicated Interconnect). One issue I found is that Google would randomly shuffle around their BGP routers which would cause BGP to flap and briefly losing all connectivity. When I raised this issue with support their answer was that this is expected behavior and we need to purchase a redundant connection. Mind you this isn’t cheap, we’re talking around $2,000 per month for a 10G connection when you add up all the GCP/colo fees. It’s pretty laughable that they can’t preserve TCP connections when they migrate their cloud routers around. I have had BGP uptimes on direct cross connects for over a year with other vendors on bare metal.
- bushbaba 3y agoIt’s about scale. Google was built for an order of magnitude greater scale where such reliability of a single link would be cost prohibitive. However in general if you need high uptime, you’ll need multiple peering links. AWS and azure also recommend the same.
- lokar 3y agoIt’s more that a single link/router/host/vm/switch/etc can never be reliable enough, so don’t waste time and money chasing that. Build your software to tolerate it. This approach is pervasive throughout all of Googles systems.
- betaby 3y agoGoogle doesn't do any 'black magic' even though they do presentations and publish papers. Their edge infra is very boring, and yes, they shuffle edge a log and their edge routers like everybody's' else - off the shelf Juniper/Cisco without any tcp session preservation.
- londons_explore 3y agoMigrating an in-use TCP session from one host to another is far from easy. I did it for a project and the number of corner cases is insane - both in the TCP protocol (what of the window has holes in? What if the connection is half closed?), but also in the OS's handling of the TCP state and interactions with userspace (will be correctly wake up a process poll()'ing a socket if we migrate the socket after a packet is received but before the kernel wakes the poll()ING process?)
- Bluecobra 3y agoI’m not saying it’s easy, but a company like Google should have no problem implementing this. If you ever used Vmotion it does a good job of migrating a live VM to another physical host. Also enterprise firewalls have no problem moving TCP/NAT state from an active to passive firewall.
- yencabulator 3y ago> Also enterprise firewalls have no problem moving TCP/NAT state from an active to passive firewall. Middlebox TCP state is a lot simpler than end host TCP state.