4 ms·
This is not really correct, and assumes the state of a cloud service (let's say load balancing) is binary. In my experience it's not. The cloud will glitch. Th
by chronid 2y ago
This is not really correct, and assumes the state of a cloud service (let's say load balancing) is binary.
In my experience it's not. The cloud will glitch. The load balancing algo will break subtly for your workload. Your traffic will get blackholed for no apparent reason. I spent a week trying to convince a cloud provider they fucked up (the time it took for them to give us someone who could run the appropriate tcpdump) once. There was no global outage.
It's on you to determine if this is important for you or not in your case, but you will need to mitigate it above a certain SLA threshold requirement. It's far for consigning CAP to history or a curiosity, which is like saying you will never have network issues if you use "serverless" stuff.
Not talking about anyone in particular, but sometimes I feel people building and using the cloud reach hubris level of surety in their systems - I worked both sides of the fence, and I know the long tails of fuckups that impact customers though...
- nijave 2y agoI remember an article--I think from Gitlab--about random latency and out of order packet issues they were seeing on GCP. Turns out it only happened under a certain amount of load since GCP's routing algorithm uses a different network if traffic is under a certain amount and switches to a more distributed mode when it crosses a threshold. In addition, sometimes you land on a bad piece of hardware but it's not bad enough to trigger the provider's monitoring. Few months ago we had a bad EC2 instance whose network would drop every 2 hours causing a bunch of random errors before recovering for a little while.