4 ms·
Honest question: why go through the hassle of multiplexing waiting in a single thread only to dispatch to a thread per client anyway? Simply using blocking IO f
by simfoo 6y ago
Honest question: why go through the hassle of multiplexing waiting in a single thread only to dispatch to a thread per client anyway? Simply using blocking IO for the clients in those threads should be much simpler right?
- deleted 6y ago[deleted]
- probably_wrong 6y agoI think the answer is in the previous article about Python linked in this article: to show that you can serve more requests with less resources if you avoid the Python/Gunicorn way and do it her way instead.
- rachelbythebay 6y agoIf you get stuck in read(), you can't do neat things like waking up when it's time to kick a client for being idle, doing other housekeeping, or cleanly shutting down the whole thing in a timely fashion. When I ^C the server, it sends the same wake condvar-poke but it twiddles the flags so the worker shuts down instead.
- pdonis 6y agoYes, using epoll with nonblocking I/O is better than blocking I/O on each worker thread. But that basically means that you are doing asynchronous programming--i.e., the exact same thing that the wonky Python/Gunicorn stack you described is doing! You're just doing it with better attention to important details. Here, to me, is the key item: The "listener" thread owns all of the file descriptors (listeners and clients both), and manages a single epoll set to watch over them. This is exactly what any async server does: it centralizes all the file descriptor management and handling in one place, and only uses workers (whether they are threads or "green threads" or whatever) to read from/write to fd's that are marked as ready in the epoll set. For your case, unless I'm misreading something, what the workers are doing in between the read/write is CPU intensive (or at least it's CPU work and not I/O work, even though it's not very "intensive" CPU work), so actual OS threads are a better choice for the workers since you can't rely on cooperative scheduling. If what the workers were doing was I/O work (for example, sending a request to a remote database and waiting for a response), "green threads" would work fine (since their only real purpose would be to organize the I/O--the actual fd's are going to be managed by the central server that manages all the fd's and checks which ones are ready for read/write). And one definitely should not try to run "green threads" for the same server in multiple O/S threads (or worse still, multiple OS processes). For an I/O bound server, one shouldn't need to anyway.
- underdeserver 6y agoBetter attention to important details is what makes or breaks a library.
- swsieber 6y agoExhibit A: Dropbox (attention to detail and executing it)
- the8472 6y ago> If you get stuck in read(), you can't do neat things like waking up when it's time to kick a client for being idle Totally possible with another thread acting as watchdog timer and sending a signal which causes the read to return with EINTR which can then check a flag whether it should retry or abort. And that's for file IO. For socket IO you can just set it to non-blocking.
- kragen 6y agoFile I/O is usually not interruptible with signals. An alternative to putting the watchdog timer in another thread is to use alarm(2) and use the kernel's built-in watchdog timer, and the default behavior for SIGALRM is probably adequate. This might be easier than non-blocking I/O.
- simfoo 6y agoGood point. Would still be possible with the threads being blocked in a read() but that would require signals and more logic in the threads, so a central multiplexing and coordinating thread seems like a cleaner solution.
- cgh 6y agoI think it's basically just the equivalent of select(2)? Once upon a time, all network servers were written this way. They were pretty fast, too.
- icedchai 6y agoMost select(2) based servers were single threaded, however. Not that this is necessarily a bad thing.
- cgh 6y agoYes, you are right, of course. I shamefully misread the article. It's closer to listen/accept/fork.
- pdonis 6y ago> It's closer to listen/accept/fork. No, it isn't, it's doing the same thing as a select(2) server would do, except it's using epoll to avoid scaling issues when you have a lot of file descriptors in the polling set. The only difference is that the workers are doing something that requires CPU, not I/O, so OS threads are being used for them (a single threaded server would be fine if the workers were just doing more I/O, like sending a request to a remote database and waiting for a response). But the worker threads are not doing any I/O management at all; they read from or write to an fd only when the central server that is calling epoll tells them to. In the listen/accept/fork model, the central server forgets about an fd once it has passed it to a handler process, and the handler process using blocking I/O.
- kragen 6y agoWhen was this time? In BSD, which introduced select(2), most network servers ran from inetd. I got a stern talking-to from the computer security folks for running a process in my .cshrc that would repeatedly finger someone at another university, where their fingerd ran from inetd, because I was making their shared VAX run unacceptably slowly. Early versions of httpd included instructions on how to run it from inetd, along with a note that you would probably regret it. Are you thinking of, like, CICS systems from the 1970s connected to SNA or something? I mean they didn't have select(2) but they did serve many clients in a single process.