6 ms·
This test doesn't include epoll in any significant fashion at all, and also doesn't apply to network operations. The libuv thing is specifically about disk file
by destructionator 3y ago
This test doesn't include epoll in any significant fashion at all, and also doesn't apply to network operations. The libuv thing is specifically about disk files, which linux doesn't really support with epoll (epoll just always considers them ready, regardless of if the requested page is in cache or not, meaning it is not helpful for the async case. by contrast, a socket without data immediately available in the kernel's buffer will not show readiness).
Since you can't tell if a disk io function would have to actually hit the disk or use the cache on older linuxes, libuv just always dispatches said operations to a worker thread pool which then calls the normal function. When it returns, it sends a message back to the other thread to indicate its completion.
If data is already present in the cache, this adds significant overhead vs just calling `read` directly - instead calling `read`, it will post a request to the worker thread queue (which may require synchronization, im not sure exactly how the impl does it tbh), wake the worker thread which calls `read`, which does not block significantly since the data is indeed already present, then it posts an event back to the main thread (which, sure, it is listening for this event on `epoll`, along with whatever else it is listening to, but that's an insignificant detail at this point) which calls the on-complete handler. All that work instead of just transferring the data from the cache to the result. Significant slowdown.
If the data is not already present in the cache, it does exactly those same steps, except this time `read` actually does block for some time, so the main thread can continue with other work in the mean time. (Of course, if the main thread has no other work to do, you've gained nothing from this!)
Hence why I said in my other comment that you ought not draw any general conclusion from this. This speedup has nothing to do with epoll vs io_uring (except maybe that linux file i/o is not really epoll compatible) and everything to do with libuv's block device implementation specifically. It is totally inapplicable to network loads entirely.
edit: in the first version I said "disk file" but it is technically block devices, of which disk files are just the most common example, but of course /dev/zero is a block device that is not a disk file.... and worth noting that /dev/zero is never going to actually load off a disk meaning it is "in the cache" all the time.
- eklitzke 3y agoIt doesn't really make sense to have an API to tell if reading from a file would block or not because page cache data can be evicted at any time. Even if there was an API that could tell you your next read would be non-blocking, there could be a race where the page cache entry could be evicted before you could do the read.
- Filligree 3y agoThat sort of race condition is only relevant if it affects correctness. Here it means a read might be unexpectedly slow, but you'll still get the same data. 99.9% of the time the data will not be evicted prior to the read, so it's still going to be a win overall.
- sgtnoodle 3y agoIf you're writing a latency sensitive program, 99.9% might as well be 0% though; you're going to end up with an asynchronous anyway.
- mort96 3y agoRemember that a filesystem read might be a disk read in the background. You don't really want to block the main thread for 2 seconds while sshfs does its thing.
- mort96 3y agoDamn, I obviously meant "network read" not a "disk read". Oh well.
- viraptor 3y agoDepends how you define that API. It could do something like "check if available, and if so, pin it until the next read (which will happen right now)". Not saying that's a good idea, but if such API was needed, it is possible.
- plorkyeran 3y agoA nonblocking read which fails if the data isn't cached and it'd have to perform io would generally be easier to get correct.
- _nalply 3y agoPlease let me try to understand you. The nonblocking read just failed. The task initiates IO then because the IO is async it lets other tasks run while waiting for data. Later, when the IO finished, the task can read the memory. Is this what you mean? I think there must be a data race somewhere or io_uring would be superfluous. Everything consisting of at least two interruptible steps at the lowest level in the CPU but where people expect no change in state is subject to some data race. It's really difficult to get this 100% correct. Something works fine for 99.99% of the time. Such a situation can be very nasty to get it right. Perhaps the step "Later, when the IO finished, the task can read the memory" cannot be done atomically.
- 0x000xca0xfe 3y agoYeah, I thought "file" refers to "file descriptor" with IP sockets included, but looks like that's not yet implemented... A couple years ago I have added io_uring to my toy webserver but it was slower than epoll. I would love to see someone using io_uring for a real networking application and report more meaningful statistics than microbenchmarks.