3 ms·
VFIO! The I/O MMU! Write your drivers in userspace, save your sanity. I realize it's hard to get a kernel that supports everything you need to accomplish this
by codemac 14y ago
VFIO!
The I/O MMU!
Write your drivers in userspace, save your sanity. I realize it's hard to get a kernel that supports everything you need to accomplish this, but it's definitely "the future" when it comes to device drivers.
However, this is a decent beginning point for people completely unfamiliar with writing a driver. The table and the makefile parts are especially important for beginners to just get freaking started.
- nitrogen 14y agoNo, don't write your drivers in user space. Especially for low-latency devices, all the user-kernel-user message passing will be problematic. You can't run at interrupt priority in user space, either.
- jrockway 14y agoMost devices are USB dongles that flash lights and shift a couple bits, not 100Gbps network cards. It follows that most people's needs will be met by a Python script and libusb, rather than a kernel module.
- lukego 14y agoI agree that kernel-space drivers are a relic of the 20th century. You can write fast network card drivers in userspace too. You map the physical DMA memory into userspace and cut out the kernel entirely. Interrupts don't even come into the picture when you're polling several packets per microsecond. Lots of those cards nowadays come with high-end packet I/O libraries that are implemented in userspace. Myricom Sniffer10G, Intel Data Plane Development Kit, ntop.org DNA, etc.
- nitrogen 14y agoI suppose. I speak from the limited experience of reverse engineering a driver for a PCI device that would freeze the entire system if its interrupts weren't acknowledged immediately.
- codemac 14y agoInterrupts are the slowest part of VFIO, but everything else is comparable. The interrupt performance can and will be fixed in the future. Polling is an option for some network uses, but interrupts are obviously technically superior for most workloads. The I/O MMU can be slightly slower than "true" DMA if done poorly, but the simple use case is actually performs the best (in my teeny experience).