5 ms·
Note that DPDK works just fine with virtio-net on GCE as well, although we have not declared official support for it. In terms of RDMA and bypass messaging, I
by jsolson 10y ago
Note that DPDK works just fine with virtio-net on GCE as well, although we have not declared official support for it.
In terms of RDMA and bypass messaging, I agree there can be substantial advantages to offloading transport layer message processing from VMs. Regrettably, I can't say anything more on the subject today.
(I lead the team that owns the virtual NICs exposed to guests in GCE)
- en4bz 10y agoDo you not use SR-IOV for GCE?
- jsolson 10y agoWe do not, although in practice as a guest you wouldn't necessarily know one way or another. Speaking from (decently informed) personal opinion here-- I see SR-IOV as a means to achieving certain networking characteristics, but one that historically didn't play particularly nicely with live migration (among other things), and one that locks you into particular silicon vendors, etc. It's a bit unfortunate that SR-IOV is often assumed as a prerequisite for high throughput, low latency, low jitter networking. On its own it doesn't guarantee any of those things (you're still beholden to the VCPU scheduler, if nothing else, and there's a lot of 'else'), and it's also not the only way to achieve those things. In particular there are a lot of ways to dedicate hardware to I/O service they aren't SR-IOV
- i336_ 10y agoWhat alternatives would you recommend to people leaning toward SR-IOV? I say this abstractly as I'm curious what your answer is and I think it might be very helpful to others considering your experience. (I just had to google "sr-iov" as I was typing this :) )
- jsolson 10y agoIn open source vhost is a pretty good option that abstracts hardware away from VMs and generally performs well. That said, SR-IOV isn't a BAD option if you don't care about live migration of VMs over a lifespan potentially longer than the underlying hardware platform. It's a good option if you can live with host-hardware-specific drivers in your guests, you're willing to standardize on a specific NIC as host hardware for the fleet where you'll be running VMs, and you're willing to deal with the resource commitment it requires to deliver on its performance promises.
- shaklee3 10y agoAre you saying that all vendors provide sriov sightly differently in an incompatible way? Because as far as I know, all the vendors have supported it for quite a while (Intel, mellanox, chelsio).
- jsolson 10y agoIt's not that their SR-IOV implementations are incompatible, it's more that SR-IOV doesn't imply anything about driver compatibility. It's is a mechanism for exposing underlying hardware (or slices of hardware) to guests and nothing more. It does not encompass register sets, queue formats, etc. So you can assign a virtual function from any SR-IOV-capable NIC into a VM using common code, but what the guest sees inside the VM will still be an Intel, Mellanox, Chelsio, etc. device. Naturally they'll need a driver for that device. If your fleet includes NICs from a variety of vendors, you'd never be able to live migrate VMs between hosts with different NIC models. Until recently, no vendor I knew of included the relevant serialization and deserialization functionality to allow for live migration at all.
- shaklee3 10y agoGood point. I didn't think of the driver incompatibility between vendors.
- kev009 10y agoIf you are trying to solve multitenancy, SmartOS makes this a non-issue by securely virtualizing Linux syscalls instead of virtualizing an entire computer like VMs.
- shaklee3 10y agoOne way you could tell is lower performance, right? Intel publishes do dpdk benchmarks, and virtio is consistently worse. Here's what I mean: http://fast.dpdk.org/doc/perf/DPDK_17_02_Intel_virtio_performance_report.pdf http://fast.dpdk.org/doc/perf/DPDK_17_02_Intel_virtio_perfor... http://fast.dpdk.org/doc/perf/DPDK_17_02_Intel_NIC_performance_report.pdf http://fast.dpdk.org/doc/perf/DPDK_17_02_Intel_NIC_performan... Sriov can do almost equivalent to the pmd,I believe.
- jsolson 10y agoUnfortunately it's a bit more complicated than that, which is why I hate that SR-IOV gets so tied up in this :/ When they say "virtio-net" there they mean virtio-net inside qemu with vhost servicing the queues on the host side (note, we don't use vhost in GCE -- our device models live in a proprietary hypervisor that's designed to play nicely with Google's production infrastructure). One could just as easily expose what looked like an Intel VF to the guest and service it in the same manner (although there are good reasons not to). One could also build a physical NIC that exposed virtual functions offering register layouts equivalent to VIRTIOs PCI BARs and used the VIRTIO queue format. If you assigned those into a guest, you'd be doing SR-IOV, but with a virtio-net NIC as seen by the guest. It also likely wouldn't perform as well as a software implementation (in its current form VIRTIO has a lot of serializing dependent loads which make it inefficient to implement over PCIe). There's some ongoing work upstream aimed at a more efficient queue format. So, yeah, "it depends" is about the best you can do. SR-IOV really just says you're taking advantage of some features of PCI that allow differentiated address spaces in the IOMMU for a single device and (on modern CPUs) interrupt routing to actively running VCPUs without requiring a hop through the host kernel. The former is handy if you want the NIC to be able to cheaply use guest physical addresses (although the IOMMU isn't free either); the latter doesn't matter if the guest is running a poll-mode driver that masks interrupts, nor does it matter if the target VCPU isn't actively running.
- shaklee3 10y agoThanks. I look forward to trying it on gce, but as others said, it would be nice if it was officially supported.
- i336_ 10y ago> I can't say anything more on the subject today. Cool. When should people start poking around GCE's network blog for Interesting Updates™ then? (Keeping in mind that I acknowledge/note that the "6 months from now" or "next year" or whatever that you say is the date you think we should start keeping an eye out, not necessarily the point in time anything will happen.)