7 ms·
An io_uring-based user-space block driver
- Joker_vD 4y agoThat reminds me quite strongly of VirtIO (block) devices... and yet the actual command format is, of course, different. Why can't we stop re-inveting things over and over?
- jmillikin 4y agoBecause different use cases require different designs? If you try to create a protocol that can work for all purposes, it'll be a poor fit for any of them and will be out-competed by more specialized alternatives. There's a reason emulators design their virtual devices to resemble real hardware (PCI, SCSI, USB) -- there's already going to be a bunch of code in the hypervisor to create fake hardware. It's also more practical to piggy-back on PCI (etc) when the spec needs to be implemented by competing vendors, since there's no kernel and no OS idioms involved. Not to mention various pre-kernel code such as EFI and bootloaders. Conversely, userspace developers really do not want to be coding up a fake PCI device with registers and interrupts and so on just to get some bytes into the kernel. They want to invoke system calls (ioctl, mmap, io_uring) and let the OS handle the details.
- Joker_vD 4y agoThe basis of virtio is literally just a ring queue, with its request descriptors looking almost exactly like structs suitable for passing to readv(2) or writev(2); the PCI shim is built on top of that and is completely optional (you can have a purely MMIO virtio device, after all). It was built this way so that KVM would not have to mimic the idiosyncrasies of real hardware: passing data as-is to the physical devices can't work for obvious reasons (even if you disregard security completely); instead in can, after minimal processing, shove it into the Linux kernel and let it take care of the rest.
- jmillikin 4y agoThe VirtIO specification for its use with MMIO[0] contains the following example device description: // EXAMPLE: virtio_block device taking 512 bytes at 0x1e000, interrupt 42. virtio_block@1e000 { compatible = "virtio,mmio"; reg = <0x1e000 0x200>; interrupts = <42>; } The next sub-section of the MMIO section is a datasheet of control registers. My point about PCI isn't strictly about PCI, it applies equally to VirtIO over MMIO. I do not ever want to have my userspace code poke at memory-mapped registers or do interrupt handling just to do the equivalent of an ioctl. --- OK, fine, maybe PCI and MMIO are irrelevant but there could be opportunities to share struct layouts. In the section describing block devices[1], there's some code listings for the request packets. A representative example is the request struct: struct virtio_blk_req { le32 type; le32 reserved; le64 sector; u8 data[][512]; u8 status; }; Take a look at that request, then look at struct ublksrv_ctrl_cmd in ublk_cmd.h[2]. There is very little the two protocols have in common. Yes, they're both doing some sort of packetized data transfer, but all of the details are different. Also, just ... just look at the size of the VirtIO specification. There is a lot there. I haven't run a `wc -l` but it would not surprise me if just the spec for VirtIO is longer than the entire patch series for ublk. Out of all that, the sum total of the virtio-blk struct layouts is something like 100, 200 lines. Is it worth going through the trouble of trying to unify these two unrelated specs just so that we can satisfy some bizarre philosophical goal of carefully avoiding new ideas? Like, if you're going to go that far, why does virtio-blk need to exist instead of continuing to emulate SCSI? Or using iSCSI for host<-> device transfer? The obvious answer is, again, because different use cases have different requirements. It's silly to cook two soups in the same bowl. [0] http://docs.oasis-open.org/virtio/virtio/v1.0/cs04/virtio-v1.0-cs04.html#x1-1090002 http://docs.oasis-open.org/virtio/virtio/v1.0/cs04/virtio-v1... [1] https://docs.oasis-open.org/virtio/virtio/v1.1/csprd01/virtio-v1.1-csprd01.html#x1-2390002 https://docs.oasis-open.org/virtio/virtio/v1.1/csprd01/virti... [2] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/uapi/linux/ublk_cmd.h?v=6.0 https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
- bullen 4y agoIs anything like this for networking done or in the works?
- jmillikin 4y agoFor networking the closest equivalent would be TUN/TAP, which lets userspace route either IP packets (TUN) or Ethernet frames (TAP).
- deleted 4y ago[deleted]
- sanxiyn 4y agoThere is XDP.
- tptacek 4y agoXDP is kind of the opposite of this, right? It's moving userland code into the kernel.
- touisteur 4y agoXDP is a lot of stuff, but I think I have someone around using af_xdp to bypass the kernel network stack and for some (and the filtering and decision of which streams, is done through some ebpf iirc) packets deliver them directly into userland buffer-queues? DPDK also has an AF_XDP backend to bridge your classical DPDK app and AF_XDP sockets.
- tptacek 4y agoAh, that's true. AF_XDP is definitely similar to userland block device offload.
- stargrazer 4y agoyes, i think there is example code where io_ring is used to get blocks into and out of XDP/kernel.
- ice3 4y agoSomehow this reminded me of this post on LKLM: - <odd>.x.x: Linus went crazy, broke absolutely _everything_, and rewrote the kernel to be a microkernel using a special message-passing version of Visual Basic. (timeframe: "we expect that he will be released from the mental institution in a decade or two"). [*] https://lkml.org/lkml/2005/3/2/247 https://lkml.org/lkml/2005/3/2/247 It's really interesting to see Linux getting more and more micro-kernel like features throughout the years.
- coolspot 4y agoNotably, the post is from a “decade or two” ago, so timeline matches.
- deleted 4y ago[deleted]
- jmillikin 4y agoI wonder if this could replace most uses of NBD (network block devices), and/or help get iSCSI into userspace where more flexible load-balancing policy can be implemented. It also reminds me of attempts to define BUSE[0][1][2], which would have been a block device equivalent of FUSE. IIRC attempts to get BUSE into the Linux kernel have been blocked for performance reasons -- the FUSE protocol isn't well designed and is only barely acceptable for VFS. If io_uring (+ careful use of zero-copy) has fixed the performance issues with userspace block devices, maybe it would be applicable to FUSE (or FUSE-v2)? I've tried using io_uring with the current FUSE protocol to reduce syscall overhead and it kinda works, but a protocol designed to operate in that mode from the beginning would be even better. [0] https://github.com/acozzette/BUSE https://github.com/acozzette/BUSE [1] https://dspace.cuni.cz/bitstream/handle/20.500.11956/148791/120397658.pdf https://dspace.cuni.cz/bitstream/handle/20.500.11956/148791/... [2] https://dl.acm.org/doi/10.1145/3456727.3463768 https://dl.acm.org/doi/10.1145/3456727.3463768
- loeg 4y agoIs BUSE significantly different from CUSE (“character device”)? https://lwn.net/Articles/308445/ https://lwn.net/Articles/308445/
- jmillikin 4y agoYep! Character devices are much closer to "stream of bytes", and from the FUSE perspective they look like a single file with limited operations (open, close, read, write). Think of something like a mouse (sending a stream of motion/click events) or a webcam (send stream of frames, receive basic control commands). If you've written even the most basic FUSE layer, you've got all the necessary handlers to implement CUSE too. Block devices operate on blocks of data identified by offset. Hard disks, CD-ROM drives, USB sticks, basically anything where it'd make sense to say "read (or write) these 1024 bytes at offset 0x10000". You can in principle implement a block device-ish API in FUSE by disabling open/close and requiring all reads/writes to be at given offsets -- IIRC this is how the "fuseblk" mode added for ntfs-3g works -- but the protocol is too chatty to be fast enough for things people want block devices for. I've also heard the kernel's block layer error handling doesn't interact well with the FUSE protocol, but I don't know the details too well on that.
- trasz 4y agoSo it’s essentially like userspace iSCSI server, but proprietary?
- loeg 4y agoWhy do you say proprietary?
- notacoward 4y agoNot proprietary, but not iSCSI-specific either. The whole idea is that you can use any protocol you like. Could be iSCSI, could be NBD, could be AoE, could be something proprietary but that's less likely than open/standard alternatives.
- trasz 4y agoIt’s obviously proprietary: it’s non-standard and specific to a single vendor. What is the whole idea, though? Serving things to kernel from userland is decades old and commonly used with both NFS and iSCSI. The fact that this particular implementation uses io_uring instead of something non-proprietary like RDMA, is just an implementation detail.
- jmillikin 4y ago> It’s obviously proprietary: it’s non-standard and specific to a > single vendor. That's not how people typically use "proprietary" when referring to open-source code developed collectively by multiple vendors, universities, and thousands of independent contributors. > What is the whole idea, though? [...] this particular implementation > uses io_uring instead of something non-proprietary like RDMA, is just > an implementation detail. When performance matters, sometimes implementation details are the whole idea. According to the patch's author at <https://lwn.net/Articles/904638/ https://lwn.net/Articles/904638/>, ublk has about twice the throughput of NBD.
- trasz 4y ago>That's not how people typically use "proprietary" when referring to open-source code Indeed, many people believe that source code being available somehow magically makes things non-proprietary. Not sure where that belief came from. An API is proprietary when it 1. Doesn't comply with existing standards, 2. Isn't interoperable, and 3. Is controlled by a single entity. >multiple vendors It comes from IBM/RedHat - a single commercial entity, not “thousands of independent contributors”. But yes, if it was a proper community project then it of course wouldn’t be proprietary. >twice the throughput Compared to NBD which can’t use RDMA at all. Again: the idea is old and not bad, it’s just that this particular implementation looks like another case of NIH.
- benlwalker 4y agoFor me, the killer use case for this is presenting logical volumes to containers. There just has not been an efficient mechanism for a local storage service in one container to serve logical volumes to another container on the same system until this. For VMs there is virtio/vfio-user, but for containers the highest performing option until this was NVMe-oF/TCP loopback. Basically, you can implement a virtual SAN for containers efficiently with this.
- lifty 4y agoVery appealing. Do you think this solution would be comparable in performance with an in-kernel storage driver?
- topspin 4y agoI'd like a built-in iSCSI volume driver for docker, podman, et al. There are third party things (netapp trident[1], etc.) but no generic driver. One would think -- given the ubiquity of SAN boxes populating racks outside of cloud operators -- you could "-v iscsi:<rfc-4173-iscsi-uri>:/mountpoint" a network block device into a container out of the box. I suppose it's difficult to deal with in cross platform way. When you read the golang source for trident you see they're just exec-ing iscsiadm on linux container hosts. [1] https://github.com/NetApp/trident https://github.com/NetApp/trident
- jimcavel888 4y ago
- sophacles 4y agoSuch as? You seem very confident about it, so why not just show the problems instead of dropping vague hints?