6 ms·
One Million Concurrent TCP connections
- imsy 15y agoGood work. On Imsy (www.imsy.com) we hope to achieve the same with Node.js on EC2. Presently the numbers are smaller (in 100K range).. but looks like it will scale smoothly till 1 million.
- henry501 15y agoIt's hard to say that something is going to scale smoothly up an order of magnitude until you're there. 900K requests is a lot of room for things to go wrong.
- nivertech 15y agoNode has memory limit of about 140K connections. I was able to accept and actually do something usable with 1M on EC2 and 3M on physical server (Erlang and Ubuntu/CentOS)
- getsat 15y agoWhat is the kernel structure overhead (in bytes) per connection on FreeBSD?
- jsr 15y agoOK, you hooked me with the title. But "FreeBSD + Erlang" was kind of a dissatisfying reason for how you achieved it. Would love to hear more details! How far we've come since http://www.kegel.com/c10k.html http://www.kegel.com/c10k.html
- getsat 15y agoThey could have done it using C, Ruby, Python, or any other language. kqueue is what makes FreeBSD (and OSX) awesome at concurrency. http://en.wikipedia.org/wiki/Kqueue http://en.wikipedia.org/wiki/Kqueue
- silentbicycle 15y agoHow does kqueue compare to epoll on Linux? I've written C code using kqueue on OpenBSD and OS X, but have only used epoll via libev (and not at especially high load). I thought the big change came from trading level- for edge-triggered nonblocking IO, but maybe the kqueue implementation is superior for sockets somehow? The main advantage Erlang has over C/Python/Ruby/etc. is that asynchronous IO is the default throughout all its libraries, and it has a novel technique for handling errors. Its asynchronous design is ultimately about fault tolerance, not raw speed. Also, it can automatically and intelligently handle a lot of asynchronous control flow that node.js makes you manage by hand (which is so 70s!). You can make event-driven asynchronous systems pretty smoothly in languages with first class coroutines/continuations (like Lua and Scheme), but most libraries aren't written with that use case in mind. Erlang's pervasive immutability also makes actual parallelism easier. With that many connections, another big issue is space usage -- keeping buffers, object overhead, etc. low per connection. Some languages fare far, far better than others here.
- deleted 15y ago[deleted]
- asomiv 15y agoYes I would say kqueue, the interface, is superior to epoll. Kqueue allows one to batch modify watcher states and to retrieve watcher states in a single system call. With epoll, you have to call a system call for every modification. Kqueue also allows one to watch for things like filesystem changes and process state changes, epoll is limited to socket/pipe I/O only. It's a shame that Linux doesn't support kqueue. But as awesome as kqueue is, OS X apparently broke it: http://pod.tst.eu/http://cvs.schmorp.de/libev/ev.pod#OS_X_AND_DARWIN_BUGS http://pod.tst.eu/http://cvs.schmorp.de/libev/ev.pod#OS_X_AN...
- gorset 15y agoI fully agree that kqueue is awesome, but what specifically is broken on OSX? I've used in extensively on that platform, and haven't run into any showstoppers.
- indygreg2 15y agoTechnical details would be interesting. Until then, here's Urban Airship's post from last year on 500k connections on Linux: http://urbanairship.com/blog/2010/09/29/linux-kernel-tuning-for-c500k/ http://urbanairship.com/blog/2010/09/29/linux-kernel-tuning-...
- chubs 15y agoThey beat me to it! I've only gotten to 500k on EC2, however i believe there's some trickery in their firewalls / NAT which is holding me back... If anyone's interested in the gory details, see: http://splinter.com.au/tag/comet http://splinter.com.au/tag/comet
- forsaken 15y agoCurious about what you've seen from the EC2 networking gear that's holding you back? Firewall not letting more connections through?
- chubs 15y agoWell, at the moment i can't pin it down to EC2, but it's the only thing i can imagine it'd be. The network/cpu/memory usage is all healthy, and there's nothing in the kernel log, so that's what i'm guessing is the cause. Although i may be wrong.
- csarva 15y agoWe've observed the bottleneck to be an upper limit on packets/sec for a given instance type. On an m1.large this is about 100k/sec. I believe it's due to the virtual NIC just not being fast enough to handle high traffic loads. The rightscale folks found the same thing: http://blog.rightscale.com/2010/04/01/benchmarking-load-balancers-in-the-cloud/ http://blog.rightscale.com/2010/04/01/benchmarking-load-bala...
- mrb 15y agoI remember reading around 2002-2004 about a sysadmin managing a very large "supernode" p2p server who was able to fine-tune its Linux kernel, and to recompile an optimized version of the p2p app (to allocate data structures as small as possible for each client) to support up to one million concurrent TCP connections. It wasn't a test system, it was a production server routinely reaching this many connection at its daily peak. If it was possible in 2002-2004, I am not impressed that it is still possible in 2011. One of the optimizations was to reduce the per-connection TCP buffers (net.ipv4.tcp_{mem,rmem,wmem}) to only allocate one physical memory page (4kB) per client, so that one million concurrent TCP connections would only need 4GB RAM. His machine had barely more than 4GB RAM (6 or 8? can't remember), which was a lot of RAM at the time. I cannot find a link to my story though...
- gtani 15y ago(Not your link but) some pointers on tuning erlang/OTP servers: http://www.erlang-consulting.com/thesis/tcp_optimisation/tcp_optimisation.html http://www.erlang-consulting.com/thesis/tcp_optimisation/tcp... http://www.trapexit.org/Building_a_Non-blocking_TCP_server_using_OTP_principles http://www.trapexit.org/Building_a_Non-blocking_TCP_server_u... http://groups.google.com/group/erlang-programming/browse_thread/thread/f47daa85ca45ed71/ http://groups.google.com/group/erlang-programming/browse_thr... (this erl mailing list thread is pretty typical, if you put up code, describe your app, hardware, network, database/external dependencies, etc, you'll get a ton of good advice about killing off bottlenecks. Another example http://groups.google.com/group/erlang-programming/browse_frm/thread/1931368998000836/ http://groups.google.com/group/erlang-programming/browse_frm...
- mmaunder 15y agoAgree with this being underwhelming. Maybe they'll hire someone who understands that Erlang is using kqueue to make this possible. Running netstat|grep like this on a high concurrency server takes a long time to run. I've never found a faster way to get real-time stats on our busy servers and would be interested if anyone else has.
- derobert 15y ago
- Rickasaurus 15y agoThis may be a dumb question (I'm not a networks guy) but how do you maintain so many connections with just 65535 ports? Can you have more than one connection per port?
- pacala 15y agoThere is one unique connection per (address, port) pair.
- Rickasaurus 15y agoAhh I see, you just give one box a ton of IPs. Thanks :)
- kornholi 15y agoNo, the pair uses the client IP which means the client can have as many connections to your server as number of ephemeral ports allowed. There is no limit on connections except ram AFAIK.
- stonemetal 15y agoNot sure why you got down voted, last time this topic came up the author was using EC2 instances for test clients, it took them 17 or so to get the number of connections to their server that they wanted. When the server IP, server port part of the 4 tuple is constant, it takes quite a few client IPs to turn 64K ports into a million.
- huhtenberg 15y agoActually, the quintet of [protocol, src_ip, src_port, dst_ip, dst_port] is what's unique (protocol being TCP or UDP)
- pacala 15y agoGood point. Each unique client (address, port) can get a separate connection on a given/fixed server (address, port).
- huhtenberg 15y agoIn absolute numbers - wow, that's impressive. It was several years ago, but I've done my share of high-concurrency stuff under Linux and the highest I got to was about 200K connections - at which point the single-threaded server bottlenecked at its disk I/O. The main issue is not the actual connection count, it's what the per-socket OS overhead is (so not exhaust non-swappable kernel memory), how many sockets are concurrently active (have an inbound or outbound data queued) and if the application can handle all the events that epoll/kqueue report. This is not a rocket science by any means, and the kernel is relatively easy to fine-tune even when the actual load is present.
- tworats 15y agoThe WhatsApp guys are very sharp ex-Yahoo guys who've had tremendous experience with scaling systems. Rick Reed is fairly legendary. Yahoo was a long time FreeBSD shop, so it's not surprising they went with that. I hope they publish how they did it - in fact let me drop them an email and see if I can convince them to do so.
- cperciva 15y agoYahoo was a long time FreeBSD shop FWIW, Yahoo still uses FreeBSD extensively.
- on4n1st 15y agoI wish they would publish details as well. Blogtweeting an achievement is meaningless unless relevant context is also given. Hardware specs? Software stack? Software tuning? sysctl tuning? We need more info, guys!
- lenn0x 15y agoIt's been done before. Here is an article using Erlang and Linux. http://www.metabrew.com/article/a-million-user-comet-application-with-mochiweb-part-1 http://www.metabrew.com/article/a-million-user-comet-applica... Part 3 is my favorite. http://www.metabrew.com/article/a-million-user-comet-application-with-mochiweb-part-3 http://www.metabrew.com/article/a-million-user-comet-applica...
- jen_h 15y agoThis article series is kind of a Bible to me. It's not going to solve all of your problems, for sure, depending on yer stack, and setting yourself up for a forkbomb isn't the wisest in all situations, but it's got a lot of good advice & is pretty good about providing you a "Okay, tweak these parameters and then try to break it" baseline.
- jen_h 15y agoA. That is AWESOME. Props, guys! B. Just how long did it take for that netstat to return? ;)
- spokengent 15y agoFWIW, netstat isn't fun to use. conntrack is better, or cat /proc/net/ip_conntrack
- Luyt 15y agoThey're on FreeBSD. conntrack is a Linux utility.
- spokengent 15y agoah ok I see...
- KonradKlause 15y agoconntrack is not netstat. I requires netfilter's connection tracking feature. Using conntrack on a 1MC system will waste even more kernel memory!
- cnlwsu 15y agoI would be curious about the hardware used, it can make it much more or less impressive. I have done a test with 1 million concurrent tcp connections using java (mina) on an Ubuntu system... but it had 64 gb of ram. It kept running for weeks under the load which I felt pretty good about.
- gabi38 15y agoHow would kqueue compare to Windows's IO completion ports in terms of performance?
- flzz 15y agoJust because you can doesn't mean you should.
- nivertech 15y agoNot many can, but some fortunate enough to actually need this stuff.
- kitsune_ 15y agoHow much ram does this single server have, what about its cores?