8 ms·
Epoll is fundamentally broken
- arielweisberg 10y agoYou probably shouldn't share epoll FDs across threads for performance reasons. A shared nothing design is likely to perform better with a simpler implementation in both the application and the kernel. I don't see sharing FDs across threads as a useful thing to aspire to. The common design I see these days is load balancing FDs across shared nothing threads. The thread that receives the notification via the selector is the thread that does the IO (no other thread has that FD). Keep adding threads as makes sense and never block let them block. A combined queue makes sense when task sizes are large. For small tasks the performance is poor. I see the queuing decision as something you make after you have already retrieved the message from the network and presented it to an application layer which makes a decision on how to dispatch. It would be cool if the kernel would do this for you under the hood! That would be amazing. Just making it correct is not enough though.
- nine_k 10y agoIf no data are shared, what is the point of having threads, as opposed to processes?
- tyingq 10y agoOne example is in NGINX. They mostly follow a "no threads" pattern, but specifically for Linux, they support "some threads" to get around issues with Linux aio. https://www.nginx.com/blog/thread-pools-boost-performance-9x/ https://www.nginx.com/blog/thread-pools-boost-performance-9x... That's not "shared nothing" threads, of course, so it doesn't answer your question at a high level. It does highlight, though, that Linux is especially challenging in this space as compared to FreeBSD. And not just because of epoll().
- arielweisberg 10y agoFD =\= data Get the data from the FD and then let the application decide the best dispatch policy.
- redahs 10y ago"Message passing" might be considered a component of a "shared nothing" system. If so, a thread-based message passing implementation might still use shared process memory behind the scenes to efficiently pass messages between local threads. Worker threads might be written so that the responsibility for managing a resource external to the process is assigned to max 1 thread simultaneously, even if the message passing abstraction uses a bit of shared process memory behind the scenes.
- BuuQu9hu 10y agoThreads are still the best way to get asynchronous filesystem I/O in a reliable way. I know that people will reply with various explanations of how to do it for some cases some of the time; however, nobody will post a way that works as reliably as threads, because none exist. At the same time, threads are a vile abomination and we should shun them whenever possible. Computers are terrible, aren't they?
- redahs 10y agoYes, in a multithreaded scenario it's possible to think of poll() and epoll() primarily as an alternative to pthread_cond_timedwait(). These functions are not about 'event dispatching' so much as they are about putting the current thread to sleep until something of interest might have occurred. Using poll() or epoll() means that the thread gets to wake up on file descriptor events in addition to intrathread signals, and epoll() seems to work fine in this regard.
- yokohummer7 10y agoSo, the kernel is free to assign a connection to a new thread even if there is a previous thread working on it? Because the kernel has no idea when the previous worker will be done processing the data. So the user has to manually unregister and register the connection each time? How about a concept like "transaction"? Maybe such as `epoll_begin` and `epoll_end`? One thread achieves exclusive access when entering the block, and releases it when leaving. Does this flaw also exist in IOCP or kqueue?
- Philipp__ 10y agoI think not, this[0] could be nice read. [0] http://people.eecs.berkeley.edu/~sangjin/2012/12/21/epoll-vs-kqueue.html http://people.eecs.berkeley.edu/~sangjin/2012/12/21/epoll-vs...
- yokohummer7 10y agoWell, just by looking at a glance, kqueue seems to pose similar problems. I mean, the problem mentioned in the OP: 1. Kernel receives some data, and wakes up A 2. A reads some data 3. Kernel receives some other data, and wakes up B (cause A didn't call `epoll_wait` yet) 4. B reads some data, leading to a race condition (out-of-order reading). kqueue's `kevent` seems to be more or less similar to `epoll_wait`, in that it doesn't seem to provide a way to notify the kernel the "we're done" signal. Am I missing here? Does `kevent` also signal the end of the exclusive access?
- Philipp__ 10y agoFrom Bryan's interview it's clear they've solved the problem like that. From article I linked, it doesn't indicate so. Take a look here[1], at 'EVFILT_SIGNAL'. But does that mean that we have to manually attach signal to monitor, and then we receive "we're done". But that's kinda similar to what you have to do with epoll, the difference is that kqueue structure encapsulates functiona of 'epoll_wait' and 'epoll_ctl'? [1] https://www.freebsd.org/cgi/man.cgi?query=kqueue&sektion=2 https://www.freebsd.org/cgi/man.cgi?query=kqueue&sektion=2 Edit: sorry for spamming with links, but I find these things really interesting, look here http://austingwalters.com/io-multiplexing/ http://austingwalters.com/io-multiplexing/ specifically at 4 steps after kqueue code, I think that explains well. So it looks like we wait for signal 'done' after which we rewatch.
- beastman82 10y agoSo what's the best way to multiplex then?
- __s 10y agoAccording to the article, use FreeBSD
- bluejekyll 10y agoThis is far too simple of an answer. There are many different reasons to use Linux and not FreeBSD. In this one case FreeBSD might be better than Linux (I offer no opinion), but there are probably other more substantial reasons why one can't easily switch from Linux to FreeBSD.
- AaronFriel 10y ago(Caveat: this is a contentious opinion on Hacker News) (Caveat #2: I'm not terribly familiar with such low level programming. Take anything I say with a grain of salt.) It's my opinion that IO Completion Ports on Windows are superior to the approach taken by *nix and BSDs. Instead of having the usermode application sleep and wake up, do some checks, etc., the application provides an entry point when an event occurs or data is available. Essentially, a callback for the kernel to use. The kernel then jumps directly to this, and can manage the threads involved, using a thread pool to balance requests. This gives much better utilization of threads than with poll/epoll/kqueue, but does place some other constraints on how the code is written. The fundamental difference is that the Unix-kin is a readiness based model. They wake up a thread to tell it that it is ready to read an event. IOCP on Windows is a completion based model, and wakes up threads with the data (or error) already present in a data structure provided to the thread.
- zzzcpan 10y ago> a completion based model, and wakes up threads with the data (or error) Which means that in this model you have to allocate and provide a buffer for that data long before the kernel is going to fill it. It's going to just sit there waiting, wasting memory. While in unix model you don't have to allocate a buffer until you know there is some data to copy from the kernel, which is easier for the user and much more efficient. Completion model makes sense if your entire networking stack lives in userspace and you can allocate memory on the lowest layer, but pass it as a reference all the way up. Or if you at least can do syscall batching, to make operating on very small buffers efficient.
- dijit 10y agoAnecdatum: I work with a guy who is pretty scary brilliant when it comes to programming; He programs on windows and his software is of a really good quality- I asked him why he used Windows and he gave two reasons: 1) Windows is mandated by HQ as the only supported desktop. You're not getting anything else to run on your desktop so you either write Windows software locally or use a VM (which is cumbersome0 2) Epoll is trash. I have successfully convinced him to make some software on freebsd because kqueue is fundamentally better. But it's shocking that epoll is so bad even compared to windows :(
- ReverseCold 10y ago> even compared to windows Windows is an expensive commercial OS, if anything it should be better, and it usually is better than the Linux desktop. Power use (no rip battery), "it just works", fast and efficient in most cases. CLI is a whole different story, but bash in ubuntu in windows is already pretty good, and there's no reason it can't be equal to bash on any other Linux distribution.
- Sacho 10y agoPowershell is pretty cool, but it's not as integrated into the "essence" of windows like sh is for linux. There's been a bunch of projects that make extensive use of it lately, like chocolatey - I hope it keeps getting better.
- ktRolster 10y agoWindows has no right to trash-talk: WaitForMultipleObjects() is one of the most depressing select-like APIs out there.
- GauntletWizard 10y agoNot being familiar, it looks pretty standard. What're it's negative attributes?
- ktRolster 10y ago
- ktRolster 10y agoThe article explains how to use epoll() correctly to solve all the problems he raises. That's not 'fundamentally broken,' because it works. The worst you can say is confusing and painful. Maybe that doesn't make good headlines, though. When you're dealing with shared resources (like a listening socket), and you are using threads, and you are trying to maximize performance with networking, confusing and painful is kind of the nature of the problem.
- zzzcpan 10y ago> The worst you can say is confusing and painful Everyone's approach towards these mechanisms is broken. Just don't treat it as a reliable notification mechanism, do your own scheduling using information from epoll/kqueue only as hints and everything will be fine.
- ktRolster 10y agoI admit, sometimes when I read these articles, I have trouble understanding why these people are having so much difficulty getting it to work. It's also hard for me to imagine a scenario where accept() takes longer than servicing a request and becomes a bottleneck. That is, why would you need multiple threads accepting on the same socket?
- majke 10y agoAuthor here. I showed an example - short lived http 1.0 connections. I haven't done benchmarks but my hunch is that you can run accept at in low tens of thousands times per second from one CPU. If you have more than, say, 10k qps the accept() might well be the bottleneck. You can imagine having a box with 64 CPU's, doing short-lived connections, and being limited by accept() done on only one CPU since it doesn't scale. Second issue is cache locality. If you do accept() in one thread only, then you will need to move the new accept-ed client socket to another worker thread. Depending on details this might not be efficient - aRFS comes into mind. (but frankly, epoll alone won't help here, you need SO_REUSEPORT with SO_INCOMING_CPU). Even if you don't agree that scaling out accept is a real concern - that's missing the point. The point is: the epoll() model should take this into account and at least support this problem. Or loudly say that scaling out accept() with epoll is not possible. But neither things happened. Up till kernel 4.5 it was impossible to do it correctly, but undocumented, from 4.5 you can use the EPOLLEXCLUSIVE flag, which I feel is a hack.
- viraptor 10y agoIs this actually the case? > Waking up "Thread B" was completely unnecessary and wastes precious resources. Epoll in level-triggered mode scales out poorly. In the analysed situation thread B was already in a wait state and there aren't enough incoming connections to immediately accept the next one. Of course the resources (CPU time) were wasted, but does that impact the performance in any way? (Assuming one "main" application on that host)
- bonzini 10y agoBecause the time wasted by all the threads trying to accept() or read() on the ready socket introduces latency for all other sockets. And since throughput is bounded by the number of threads divided by inverse of latency, increasing the latency makes you lose in scalability. In addition, because a single "readiness event" has to wake up many threads, it can introduce lock contention and cacheline bouncing in the kernel's data structures.
- lngnmn 10y agoIt is pthreads which are fundamentally broken indeed, not epoll.
- unscaled 10y agoAll of the problems outlined are relevant whether you pre-fork into multiple processes or use different threads. I don't see what it has to do with pthreads.
- bluejekyll 10y agoI don't feel like this was given enough time: > One option is to use SO_REUSEPORT and create multiple listen sockets sharing the same port number. This approach has problems though - when one of the file descriptors is closed, the sockets already waiting in the accept queue will be dropped Yes, it's a problem if you use it across processes... but if you have a single long lived process with each thread listening on a separate queue, isn't that the simplest solution?
- richardwhiuk 10y agoThat process isn't allowed to fork, I assume, which is problematic.
- bluejekyll 10y agoI'm not proposing a fork either. Each thread would open a separate socket with the SO_REUSEPORT set. The kernel would then have a separate queue per socket. Please correct me if I'm wrong.
- richardwhiuk 10y agoI'm saying that none of those threads can start any programs. That may or not be required by your use-case, but it's certainly a disadvantage.
- xorblurb 10y agoI'm curious if you can really design a practical API that avoid all the issues the author talk about. Even on something as simple as interrupt delivery to a single consumer, you MUST be prepared to handle merged and spurious interrupts -- I would argue that any driver that is not prepared to do so (in the general case) is buggy. Maybe it's easier to do perfectly with an epoll/kqueue API for some reasons, but, without having tried to think much about it, I can't imagine why it should be. I have the intuition this is way harder. Actually I'm not even sure if I can have any intuition about the difficulty to achieve the behavior wanted by the author of that article, because the author did not actually specified the behavior he desires...
- bogomipz 10y agoI have a question. The post states: >"This is because "level triggered" (aka: normal) epoll inherits the "thundering herd" semantics from select(). Without special flags, in level-triggered mode, all the workers will be woken up on each and every new connection." Isn't this behavior similar to disk I/O where the kernel wakes up(via wake_up()?) all tasks that are sleeping on a wait queue for disk I/O? I believe all processes sleeping on a disk I/O wait queue will be woken up regardless of whether their disk I/O is complete or not and will be put back to sleep in the whre it is not.
- rdtsc 10y ago> The best and the only scalable approach is to use recent Kernel 4.5+ and use level-triggered events with EPOLLEXCLUSIVE flag. This will ensure only one thread is woken for an event, avoid "thundering herd" issue and scale properly across multiple CPU's Exactly use EPOLLEXCLUSIVE for accept, that seems to work and has been in the kernel for more than a year. For reads, how reasonable is to even share the file descriptor with multiple threads. Just hand it to a thread and let that thread (possibly bound to a CPU core for possible more performance boost) handle it from then on. Am I crazy to be surprised by that architecture choice?
- bitwize 10y agoThe least broken modern OS when it comes to async I/O is Windows. I/O completion ports are the correct solution when it comes to waking up one of multiple threads to handle I/O events.