3 ms·
The thing I learned that blew me away is that with their new networking stack, the cross-AZ communication within a region is "always <2ms latency and usually <1
by pixelmonkey 12y ago
The thing I learned that blew me away is that with their new networking stack, the cross-AZ communication within a region is "always <2ms latency and usually <1ms".
As mentioned in the slides, that's the same latency usually associated with SSDs and 100X better than the latency typically associated with cross-region networking. This suggests to me that running distributed databases in multiple AZ's has almost no latency penalty (e.g. you won't even paying eventual consistency / replication lag taxes that you might think are a danger). That's pretty damn cool.
- spullara 12y agoWe see this using FoundationDB across 3 AZs. Absolutely not a problem at all with great performance and reliability.
- nitinics 12y agoIndeed. It would be interesting to know the distance between two AZs, provided 2ms is the RTT. And considering many factors including - switching and network latency including encryption / decryption on cross AZ traffic (if they do)
- thecabinet 12y ago> The above shows the US East region in Ashburn, Virginia. It has five availability zones, and these are protected areas that are isolated from each other by a couple of kilometers ...
- runeks 12y agoI wonder what it would take to get it further towards the speed-of-light limit. 2 ms for, say, 3 km is 1500 km/s, or around 0.75% the speed of light in fiber (~200,000 km/s). I wonder what the bottleneck is.
- jpgvm 12y agoI was a big Xen guy back in the day. The big takeaway from Hamilton's talk was actually something that dawned on me when they released the c3/m3/i2 instances which is what they have done with SR-IOV. The software stack that implements the AWS VPC fast path must be running on the NIC itself. This is a big deal. Being able to implement the require isolation/tunneling/encapsulation logic and routing in the NIC is a huge performance boost and drastically simplifies the hypervisor. Well, the software portion of the hypervisor atleast. If you listened closely to the talk he laid out the differences in latency impact of each layer in the networking stack, from the fibre (nanoseconds) up to the software stack at some number of milliseconds, several orders of magnitude worse than anything else in the stack. The reason why the quoted latency is so large is how para-virtualized networking is implemented. In order for it to be (bandwith) performant it needs to use large ring buffers with many segments. This is very throughput optimised and latency suffers a ton as a result. By moving all of the queuing to the NIC you get a bunch of benefits, namely the dom0 (basically the host in Xen terminology) is no longer involved in pushing packets and you are not incuring the cost of the Linux networking stack 2x. In the paravirt model the skbs are transferred across the circular buffer into the host OS where they are injected into a virtual interface, thus traversing the full net stack again. In the SR-IOV model the address space of the virtual NIC is mapped into the guest OS using Intel's IO-MMU extensions and the guest is then able to communicate directly with the NIC, thus 100% bare metal performance. If SR-IOV was the only improvement it would be impressive, however it's the consequence of it's existence which makes the biggest difference. If the guest is talking directly to the NIC, then all of the encap/decap is HW accelerated too and in theory this means the full networking stack is end-to-end in HW.
- Andys 12y agoNote: standard out-of-the-box SR-IOV allows VLAN tagging/stripping outside of the guest control, maybe they are simply using this in conjunction with a layer 3 switch to handle the VPC stuff?
- electrum 12y agoAmazon VPC is far more complex: https://www.youtube.com/watch?v=Zd5hsL-JNY4 https://www.youtube.com/watch?v=Zd5hsL-JNY4 (excellent talk from one of the creators of VPC)
- nandemo 12y ago> This suggests to me that running distributed databases in multiple AZ's has almost no latency penalty (e.g. you won't even paying eventual consistency / replication lag taxes that you might think are a danger). Yes, Hamilton points that out in his talk: https://www.youtube.com/watch?v=JIQETrFC_SQ https://www.youtube.com/watch?v=JIQETrFC_SQ