3 ms·
At around 4 GHz and 2 IPC, it's about 1000 CPU instructions per I/O. Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace
by devit 5y ago
At around 4 GHz and 2 IPC, it's about 1000 CPU instructions per I/O.
Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace, the only thing that is really needed is to write the operation type, size, disk position and memory position to the ring buffer, update the buffer position and check for flush, doable in around 8 CISC instructions (plus the slowpath), so around 100x inefficient.
Without the direct mapped ring buffer and with a filesystem, you need a kernel to translate from uring to the hardware ring buffer, and here it still seems around 10x inefficient as around 100 instructions should be enough to do the translation (assuming pages already mapped in the IOMMU, that you have the file block map in cache, and that the whole system is architected to maximize the efficiency of this operation).
- throwawaylinux 5y ago> Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace, the only thing that is really needed is to write the operation type, size, disk position and memory position to the ring buffer, update the buffer position and check for flush, doable in around 8 CISC instructions (plus the slowpath), so around 100x inefficient. This type of thing is not measured in instructions anymore, it's measured in off-board operations and latencies. Cache misses, latencies to poke MMIO and get back interrupts if necessary, DMA transfer to complete, device access time. In this case it seems the hardware is theoretically capable of about 12M so the core mostly be just waiting for that.
- Flow 5y agoWill DMA from external devices pass L3 or some other CPU cache? Or are all accesses to the freshy DMAed-to memory be cold