10 ms·
Linux 3.9 introduced a new way of writing socket servers
- buster 13y agoCool, i didn't know that would be possible with sockets, sounds like a nice option. Although i wonder how efficient it is, but it may be worthwhile to spawn [number of cores] x [node.js | python | ruby] servers which themselves only run asynchronous functions, greenlets, etc. in a single thread..
- audidude 13y agoThis could be useful for periodic tracing/profiling as well. Simply have a second instance with all debugging symbols and tracing enabled, but only accept() a client every X seconds.
- lttlrck 13y agoThat's a really neat idea. Thanks, it could be useful.
- hosay123 13y agoSadly it doesn't work like that.. if 2 processes have the same port number bound, then approximately 50% of clients will hash onto the second receive queue. If the debug process only accepts a few connections every so often, then nearly 50% of traffic will essentially be dropped on the floor It's also not possible to occasionally listen and unlisten.. that causes the hash modulus to change, sending traffic to the wrong sockets and (most likely) resetting all existing connections
- pritambaral 13y agoThe hash modulus reset issue is being worked on. Source: the original lwn posting.
- deleted 13y ago[deleted]
- haberman 13y agoInteresting. It seems like one potential hazard is that bonafide port conflicts are not detected. If SO_REUSEPORT is preferred for performance reasons, and most/all servers are using it, then starting up a server that uses the same port as an existing service becomes a silent error. It could even work as expected for a while (since the kernel gets to arbitrarily decide what port to deliver incoming requests to) only to intermittently fail later.
- geocar 13y agoI can't imagine people will start using SO_REUSEPORT by default, since the "performance reasons" are a happy accident of having a hint (that the process wants wakeups distributed across all CPUs). I'd rather get that hint in another way- perhaps by sharing an epollfd with multiple processes. I would however like SO_REUSEPORT to run experiments: Right now we use iptables/tc to direct some traffic at "new versions" of some of our systems so we can run tests with live data, but connection tracking for localhost is lame. I'd much rather use SO_REUSEPORT.
- wicknicks 13y agoI imagine production servers would run monitoring processes to remove such "listening bugs".
- MertsA 13y agoThe second process would have to be run under the same user for that to happen though so a real production system would probably never be impacted by this but what would be nice is a flag to restrict the port to children of a particular PID or just lock it to one particular PPID.
- ikeepforgetting 13y agoListening to the same port requires processes with the same uid.
- rbanffy 13y ago> then starting up a server that uses the same port as an existing service becomes a silent error. Only if it has the same uid as the other one. It'd also be trivial to check whether the other processes listening to your port are "friendly" (as in "you don't want both Apache and Nginx listening on port 80").
- zzzcpan 13y agoMeh. SO_REUSEPORT doesn't change the way socket servers are written. I was expecting something, like syscall batching for sockets, but not this.
- rdtsc 13y agoThey way I understood it, it does change because it simplifies the server. There is no need for one top level listening process/thread. Each separate process/thread can open the listening socket independently. (Btw, there is another interesting forking-for-client-connection pattern in Erlang. Instead of forking off and handling the client connection in a separate process, instead handle the client connection in the accepting process but fork-off another process to continue accepting. In general, just a process pool, that should be easier to set up with this new feature).
- RivieraKid 13y agoPretty much, it's just some small unimportant technical detail.
- rdtsc 13y ago> Now the question is why to bother with multiprocess socket servers at all - aren't threads and events better? There's at least one good niche for them - dynamic languages like Python or Ruby, which need multiple OS processes to achieve real concurrency [my emphasis] That is not true. It is an often repeated misconception. It makes it sound like Python creators were just incompetent and just stuck threads in there even though they are completely useless. In fact Python's threads work well for IO concurrency. I used them and saw great speedup when accepting and handling simultaneous socket connections. Yes you won't get CPU concurrency, but if your server is not CPU bound you might not notice much of a difference. IO concurrency is real concurrency. In 8 years using Python for fun and professionally I probably wrote more IO concurrent code than CPU concurrent code. Even then for CPU concurrent code I would have had to drop into C using an extension (and there you can release the GIL anyway). Now, the obvious follow up is that in case of IO concurrency you are often better of using gevent or eventlet. You get lighter weight threads (memory wise) and less chances of synchronizations bugs (since greenlet based green threads will switch only on IO concurrency points, socket reads, sleep and explicit waits on green semaphores and locks).
- sluukkonen 13y agoWhat he means is true parallelism (although alternate ruby and python implementations have it).
- rdtsc 13y agoDownloading 2 web pages at the same time without one blocking another from completing is true parallelism. The request is sent for one, while it is in progress (maybe server is slow), another one can go out and come back with data. This can happen for hundreds or thousands of them. These are executed in parallel. So we got concurrent units of work executed at the same time, I fail to see how that is not parallel. Now this is IO concurrency but it is real concurrency. Adding CPU concurrency would be very nice. It might speed things up a bit, or it might not. It really depends. As an example consider haproxy. The little proxy that could. It handles large amounts of concurrent connection in parallel and it is single threaded in its default configuration. I've heard of 100k connections. It deals with IO concurrency. Chances are, making it multi-threaded might not dramatically improve its performance (it might even slow it down).
- bborud 13y agoWhy does the blog posting only mention fork and prefork as options? A very common way to design servers is to do multiplexing IO. The one-connection-per-thread/process isn't the only way. That being said, this option can simplify things -- removing the necessity of having some moving part to distribute connections across completely independent processes.
- joosters 13y agoThey aren't mutually exclusive. You can have multiple processes performing non-blocking I/O, as a way of scaling over several cores without multithreading.
- bborud 13y agoExactly.
- jerf 13y ago"Why does the blog posting only mention fork and prefork as options?" Because this is a Linux kernel feature involving sharing a socket amongst multiple OS processes, and is therefore only interesting to talk about if you are using multiple OS processes. It's not a generalized primer on all techniques of handling IO.
- joosters 13y agoYou never needed to prefork. One process can open a listening socket and share it with an unrelated process via file-descriptor passing.
- rdtsc 13y agoThat is a cool trick! Does that work via pipes and maybe also unix sockets? I suspect latency might still be slightly better with a pool of pre-forked processes/threads.
- bratsche 13y agoYes, you can do it with unix sockets.
- joosters 13y agoYes, you can do it with unix sockets. Not sure about the latency, but you can pass the listen sockets in advance, rather than having one process accept()ing incoming connections and then passing those to other processes to handle. So unless you're binding to new ports all the time, it's all just a little extra startup work and won't impact the performance of the server.
- bdarnell 13y agoIt uses unix sockets. Latency is the same as the standard pre-forking model; the only difference is that file-descriptor passing lets you manage the worker processes independently instead of requiring them to have a common parent process (this is important when rolling out new code to a service with a lot of active connections, since it's disruptive to restart all the workers at once). Here's a demo in Python: https://gist.github.com/bdarnell/1073945 https://gist.github.com/bdarnell/1073945
- rdtsc 13y agoThat is pretty neat, thanks for sharing. I added it to my notes and code snippet for future.
- 13y ago
- gargoiler00 13y agowhy would anyone still be using threads or processes these days? :/ hardly scalable or efficient.
- lttlrck 13y agoTo take advantage of multiple cores?
- gargoiler00 13y agoYeah because I have as many cores as I have concurrent HTTP requests, and obviously it's CPU bound...
- erichurkman 13y agoSockets are used for a lot more than just serving HTTP requests.
- gargoiler00 13y agoThey certainly are. But my point is that if you're only servicing the same number of IO connections as you have cores, it's not extremely scalable. Most networking servers should be dealing with hundreds or thousands of concurrent connections.
- bdarnell 13y agoThe linked article didn't make this clear, but this feature is mainly designed for process-per-core models, not process-per-connection. The problem you run into with most existing process-per-core systems is that you can't ensure an even distribution of load across the processes without introducing extra overhead. SO_REUSEPORT offers some convenience when changing the number of processes, but the real benefit is that in this mode the kernel uses a better load-balancing scheme.
- jerf 13y ago
- cperciva 13y agoFor what it's worth, BSD has had SO_REUSEPORT since BSD 4.4-Lite.
- nemetroid 13y agoFor anyone curious: released in 1994.
- sounds 13y agoMore information from 2010 about the way to do that in Linux: http://stackoverflow.com/questions/3261965/so-reuseport-on-linux http://stackoverflow.com/questions/3261965/so-reuseport-on-l... This is an example of a major downfall with free software: a developer decides he needs a feature so he implements it without taking any effort to see what has been done before – and more importantly, why. It leads to the project sprouting thousands of new features while nothing achieves the polish and completeness of the original idea because the developer moved on to something newer and shinier. I can't find the original blog post where I read the idea, but I did find one on Coding Horror: http://www.codinghorror.com/blog/2008/01/the-magpie-developer.html http://www.codinghorror.com/blog/2008/01/the-magpie-develope... The Linux kernel solves this by having Linus, who has the long term perspective and the commitment to keep the project moving forward. I'm not claiming he's perfect, just that having him is the correct solution to the problem. Obviously here is someone who thinks the 3.9 kernel has a new feature he needs all the while ignoring past socket work.
- simonw 13y ago"This is an example of a major downfall with free software: a developer decides he needs a feature so he implements it without taking any effort to see what has been done before – and more importantly, why." Reinventing the wheel is certainly a common flaw of developers, but I don't see what it has to do with free software. Are you suggesting that it's less present in non-open-source software development?
- sounds 13y ago
- fooyc 13y agoThis is likely to consume more memory, because of copy on write pages (or lack of thereof). Implementing the prefork model by spawning unrelated processes (by opposition to forking from a common parent process) is likely to consume more memory: each process is unrelated, and do not share copy on write memory pages with other processes.
- nullc 13y agoShared library code already gets shared, so this may not be as bad as you think.
- justincormack 13y agoYou could use SO_REUSEPORT with threads too. Linux threads are processes after all, with tweaked clone() options.
- rgarcia 13y agoThis seems very relevant to people using Node, considering it has basically standardized around the "pre-fork" [0] model as a way to use more than one core. It'll be interesting to see where this goes. [0] http://nodejs.org/api/cluster.html http://nodejs.org/api/cluster.html
- gwu78 13y agoThis is the -T option in W.R. Stevens' sock utility. See Appendix C to his December 15, 1993 book on TCP/IP. 1993.
- MalcolmEvershed 13y agoIt seems like this could help solve the thundering herd problem [0][1][2], no? [0] http://en.wikipedia.org/wiki/Thundering_herd_problem http://en.wikipedia.org/wiki/Thundering_herd_problem [1] http://stackoverflow.com/questions/15636319/why-is-accept-mutex-on-as-default-in-nginx http://stackoverflow.com/questions/15636319/why-is-accept-mu... [2] http://uwsgi-docs.readthedocs.org/en/latest/articles/SerializingAccept.html http://uwsgi-docs.readthedocs.org/en/latest/articles/Seriali...
- wmf 13y agoProblem was already solved: "In modern times, the vast majority of UNIX systems have evolved, and now the kernel ensures (more or less) only one process/thread is woken up on a connection event."
- MalcolmEvershed 13y agoI believe that quote from [2] is referring to simply calling accept(), but modern socket servers use epoll() (or similar) before accept() which I think still has the problem (because I've run strace on nginx and uwsgi and I'm pretty sure I saw all processes wake-up from epoll()). So I'm thinking that with SO_REUSEPORT, each server process would have a different socket to epoll() on, and the kernel would only wake-up one process on a new connection, thus, solving the thundering herd problem for modern servers.
- mrottenkolber 13y agoThe same model I use in my soon to be released web server. :) Have a thread pool compete for an accept-lock. Performance isn't that bad actually. About the same as thttpd.
- Refefer 13y agoI'm a bit more worried about the security aspect of it. Let's say that we are running a server on a port which uses this option to allow multiple processes to bind to it. What's to prevent a rogue process, perhaps with malicious intent, from starting up and siphoning off requests willy nilly? Sounds like a great way to implement a hard to detect MITM attack. What would be nicer, I think, is if socket reusing was bound not only to the same uid but also to the process listening to it.
- pfraze 13y agoYou can mitigate that risk by using one of the first 1024 ports, since they require root access.
- takeda64 13y agoAs I understand you need to have the same EUID to be able to bind to the same port.
- subim 13y agoThat's right. This article doesn't mention it, but the LWN article it cited (https://lwn.net/Articles/542629/ https://lwn.net/Articles/542629/) does.
- halayli 13y agonginx already scales by spawning multiple processes. The worker processes share the listening file descriptors from the parent master process which allows the workers to accept connections on the listening fds.
- robbles 13y agoOne detail that doesn't seem to be mentioned here or in the linked article is how the multiplexing of sockets is actually handled at the kernel level. Does the kernel use some sort of round-robin approach to assigning client sockets to processes waiting on accept()? This is one area where I'd imagine a dedicated master process would be beneficial, as it could implement "smarter" load balancing based on the health and response times of its child processes.
- jkn 13y agoAm I right that this makes it trivial to deploy a new version of my server with zero downtime? I can just start the new server to handle new connections and tell the old one to stop accepting connections and quit when existing requests are completed, no need for another layer routing?
- pfraze 13y agoThat seems correct to me.
- DonPellegrino 13y agoThat's exactly how I do it for my Node programs. Any service that I want to have 100% is using the cluster module, resulting in multiple processes listening to the same port. When I want to update, I replace the files and kill the processes one by one.
- caf 13y agoYou could already do this by having a way for the new version to connect an AF_UNIX socket to the old version and request that the listening file descriptor be passed from old to new.
- Amadou 13y agoIs SO_REUSEPORT really all that much better than a server process that hands off incoming connections to other independent processes via an AF_UNIX socket with sendmsg/recvmsg? If I understand SO_REUSEPORT right you let the kernel decide everything - access control, receiving process, timing, etc in exchange for not having your own process doing the same thing. Since that simplistic approach is the kind of thing that can be implemented in about 100 lines of user-space code doing file-descriptor sharing with sendmsg/recvmg via AF_UNIX sockets, I don't see the benefit of pushing that complexity into the kernel. Especially since if you want to exercise any greater level of control you'll just have to roll your own AF_UNIX based code anyway.
- fexl 13y ago"in the fork model a number of processes can grow uncontrollably." You can use setrlimit to prevent that. Plus, your application is likely to have direct control over forking anyway.
- IgorPartola 13y agoSo this is inereating, except in the real world your parent process does more than the article implies. The big thing it is in charge of (and the thing that I have seen many of them get wrong) is (a) keeping the child processes running/restating them when they fail and (2) performing graceful config or code reload. The OS has no business doing the latter and would have a very hard time doing the former. In fact I have seen issues where gunicorn failed miserably simply because it did not handle a bad import in a child process. Tornado as of the latest version I had used (2.0 I think) did not have any ability to check for dead child processes. I am sure there are more examples of this done wrong than right. This is an interesting option for several use cases but you still need a parent process to monitor things. Perhaps at some point upstart or systemd will get good enough to monitor multiple processes per daemon in real time. Until then, meh. Edit: actually, one cool thing you can do with this is code reloading. You simply have your parent process start more workers that attach to the same socket, then kill the old ones. That way the idea of code or config reloading doesn't need to be baked into every part of the worker.
- sorbits 13y ago> in the real world your parent process does more than the article implies […] keeping the child processes running/restating […] performing graceful config or code reload […] The article suggests you let http://supervisord.org/ http://supervisord.org/ (or similar) take care of these things.