5 ms·
It's fun to speculate about how other clouds do things :) > There's absolutely no way that they would get the performance I'm seeing from an emulated disk. We'
by jsolson 9y ago
It's fun to speculate about how other clouds do things :)
> There's absolutely no way that they would get the performance I'm seeing from an emulated disk. We're talking to real hardware, exposed via PCI passthrough.
There's a wide spectrum between "emulated" and "real hardware, exposed via PCI passthrough". Passing through to PCI hardware, in and of itself, gains you very little in terms of absolute guest-visible performance versus eliding all VMEXITs via other means, but it has other important characteristics that I suspect AWS very much wants in the c5 family.
> I would assume it's something like "NVME interface hardware" + "ARM CPU which implements the EBS protocol" + "25 GbE PHY", but that guess is based solely on "that's how I would design it".
I would expect something along these lines, although I'd be a little surprised if they bothered putting the NVMe bits down in silicon. My personal guess would be a good silicon DMA engine and PCIe interface married to sufficient general-purpose processing (ARM SoC, FPGA, etc.) to keep pace with the NVMe Command/Completion queues.
Regardless of how they've implemented it, the end-to-end result seems to hang together very nicely. Kudos to the team at AWS.
(note: I work on Google Compute Engine's hypervisor; my speculation about AWS really is speculation about how they'd do this — the two companies have very different engineering approaches, so I may be entirely wrong trying to project onto theirs — like I said at the top, it's fun to try :)
- kijiki 9y ago> I would expect something along these lines, although I'd be a little surprised if they bothered putting the NVMe bits down in silicon. You and Colin both know they bought Annapurna Labs, right? We don't have to speculate _that_ much about what is probably going on here...
- cperciva 9y agoThat's why I said ARM for the EBS protocol handling rather than MIPS. :-) I guessed a hardware NVMe interface because that seems like something which could be acquired more or less off-the-shelf, thus minimizing the engineering risks.
- jsolson 9y agoI do, hence my speculation about putting the NVMe in firmware instead of hardware :) (Colin's speculation in a peer reply is also reasonable -- personally I've seen enough errata in "off the shelf" IP to shudder at the idea of anything in silicon that doesn't have to be, but fundamentally I'm a SWE, so that would be my take, wouldn't it)
- cperciva 9y agoThe way I see it, everything has errata... but if you're taking something off the shelf, it's more likely that someone else already found them. :-)
- _msw_ 9y agoAnd I could speculate on hypervisor bypass in Andromeda 2.1 ;-)
- jsolson 9y agoIndeed! I didn't leave too much to the imagination with my replies on the original post[0], though. Honestly, I'm mostly curious about how much of "KVM" you're running that's stock, how much is modified, and how much of the userland is running on the far side of PCIe rather than in host ring3 (particularly given "C5 instances are built using a new light-weight hypervisor, which provides practically all of the compute and memory resources to customers’ instances."). [0]: Especially this one: https://news.ycombinator.com/item?id=15641391 https://news.ycombinator.com/item?id=15641391