9 ms·
Achieving 100k connections per second with Elixir
- dclusin 8y agoWould be helpful to know the hardware/instance size they used for these tests. TFA doesn't explicitly state it.
- jschniper 8y agoThey mention in the Ranch section that "In this test, we set it to 36—the number of CPU cores on our c5.9xlarge."
- zambal 8y agoWe used Ubuntu 18.04 with the 4.15.0-1031-aws kernel, with sysctld overrides seen in our /etc/sysctl.d/10-dummy.conf. We used Erlang 21.2.6-1 on a 36-core c5.9xlarge instance. To run this test, we used Stressgrid with twenty c5.xlarge generators.
- lstodd 8y agoomg. 100K/sec was achieved by yours truly 10 years ago on a contemporary xeon with nothing but nginx and python2.6 - gevent patched to not copy the stack, just switch it. (EDIT: and also a FIFO I/O scheduler) Why does this require 36 cores today??
- jacobn 8y agoWas your benchmark for requests/sec or connections/sec?
- lstodd 8y agoSingle-request connections. Response required consulting memcached and updating it from postgres if out of luck, which was very rare but still needed (and patching then-existing postgres C client to be async aware was an undertaking)
- jasonlotito 8y ago> Single-request connections. What does that mean? You keep qualifying "connections." It's a connection. It holds onto it's connection for X period of time. An HTTP request is just a single-request connection, which is NOT what this article is discussing.
- lstodd 8y agoOne HTTP connection, one request, one response, connection closed. I admit I didn't first see that they actually don't do any i/o over those connections. Well, you know, handling x accepts() per second and holding onto y fds is even less than nothing to be proud of.
- jasonlotito 8y agoSo yeah, those are generally considered to be requests per second. Apples and oranges.
- lpgauth 8y agoDuh. Of course a C event loop while be faster at accepting connections, that's not the point of the article.
- lstodd 8y agoThey boast only accepting 100K connections per second, not pushing back a meaningful response? Why this is even here then?
- rozap 8y agoIf you read the article, in the third or so paragraph. > What this means, performance-wise, is that measuring requests per second gets a lot more attention than connections per second. Usually, the latter can be one or two orders of magnitude lower than the former. Correspondingly, benchmarks use long-living connections to simulate multiple requests from the same device.
- lstodd 8y agoYour point being? I was talking of single-request connections.
- jasonlotito 8y ago> I was talking of single-request connection. Yes. Which is not what's being discussed here.
- dzik 8y agoWould you mind sharing the details? (URL maybe) I think limiting factor might be not number of cores and outside of erl scope, that is eth card they used, network infrastructure, etc. Even Elixir could be something that impacts the tests.
- lstodd 8y agoThere is no url summing the details unfortunately. The work in some unknown state is at https://code.google.com/archive/p/coev/ https://code.google.com/archive/p/coev/ Without the business logic (which was in django IIRC) and deployment details, obviously. Very outdated and some later patches might be missing. No one was interested, you see. I'd be surprised if there were problems with network, and if there were, that should have been obvious in the metrics. Maybe the metrics were inadequate
- dzik 8y agoSorry, where do the authors claim they achieved >100k connections per second?
- lstodd 8y agoI'm the author, and that's the truth. Can't see how this can be replicated as a controlled experiment nowadays, unfortunately. But if you define exactly what's a request, what's a response, and what the connection/response ratio is let's have a race. Like, you set the parameters, and whoever serves that on lower-capability hardware wins. Py3 plus low-level C/Rust hacks vs Elixir, say.
- dzik 8y agoThat's the thing. You can always hack something in C to prove there is a better way for a specific task. In the past I did things like that just for fun. But in the real world it does not work like that. You buy into things as a whole, accepting their pros and cons as a whole. If you need to hack - change your tools.
- benfolred 8y agoYou are comparing apples and oranges. They are purposely holding the connections around for 1+10%seconds. So first of all, it means that, for a rate of 100k conn/s, they are going to have around 200k open connections after a second. This already imposes a different profile than 100k single request connections per second. You are also assuming that they need 36 cores to achieve 100k connections per second, which is likely not the case since they quickly moved the bottleneck to the OS. I am assuming they have other requirements that force them to run on such a large machine and they want to make sure they are not running into any single-core bottlenecks (and having a large amount of cores makes it much easier to spot those).
- Thaxll 8y agoI highly doubt you were able to do 100k connections/sec 10 years ago with the same hardware, you must be confused between requests/sec and connections/sec very different things.
- StreamBright 8y agoNothing tells more about an engineer than the last undocumented unreproducible hello world micro benchmark conducted by her once and only once some years ago that beats a real world application in terms of req/s leaving out latency profile.
- dclusin 8y agoskimming fail :(
- dzik 8y agoDid starting more acceptors than the number of cores make any difference?
- dzik 8y agoThis article is quite good, especially part about bottleneck caused by single supervisor in ranch. However I have to say that title is a bit misleading because all of this has nothing to do with Elixir, it's all about Linux kernel and Erlang, cowboy and ranch are written in Erlang. Having said that, I will add that I think it is good to have Elixir.
- dnautics 8y agoPresumably they were using cowboy through Elixir. It's not hard, the module is just called :cowboy instead of cowboy.
- dqv 8y agoOr `Cowboy` with `alias :cowboy, as: Cowboy` ;)
- latch 8y agoStick with the :module notation. This makes it clear that you need to be in "erlang mode"...indexes starting at 1, [probably] charlists instead of binaries, and possibly weird (from an elixir programmer's point of view) argument order.
- dnautics 8y agoTo be fair, it's a rare thing to be using an index in elixir at all.
- csisnett 8y ago"is a bit misleading because all of this has nothing to do with Elixir" Stressgrid is written in elixir though, https://gitlab.com/stressgrid/stressgrid https://gitlab.com/stressgrid/stressgrid
- dzik 8y agoPoint taken and I am already looking at stressgrid, "millions of users" is definitely a selling point to me. It is actually quite hard to generate enough and correct traffic to stress test large distributed systems.
- deleted 8y ago[deleted]
- dnekencjfkerf 8y ago> What this means, performance-wise, is that measuring requests per second gets a lot more attention than connections per second. Usually, the latter can be one or two orders of magnitude lower than the former. does anyone know how does 100k connections compare with other servers?
- Thaxll 8y agoIt's probably easy to do with Java / C# and Go, they're using a 36 cores machine to achieve that with fast CPU, meaning that you need 3000conn/sec per core, very doable with recent frameworks.
- ralusek 8y agoShould be possible just fine with NodeJS, so long as it's clustered to run an instance per core. The order of magnitude(s) differentiator for server performance really comes down to whether or not the architecture is blocking or non-blocking.
- holoduke 8y agoWe run about 20k connections per second with nodejs on a 12 core machine. All node is doing is parsing cached JSON, modifies it and serve it back to the client. One server has an uptime of 560days without any memory/performance issues.
- ioquatix 8y agoOn my desktop computer with a single thread, Ruby can handle about 2000/conn/s. I'm just going to check a single thread with a similar C++ implementation.
- deleted 8y ago[deleted]
- adontz 8y agoWith Python/uvloop I can easily get 10K-12K connections per second per core, so 36 cores will be fine with Python too.
- muststopmyths 8y ago>Finally, the connections per second rate reaches 99k, with network latency and available CPU resources contributing to the next bottleneck. Can someone educate me on what they might talking about here ? CPU is ~45% in their final graph. I don't know what network latency means in this context though. Roundtrip for a TCP handshake ? That seems unlikely.
- dzik 8y agoIt means CPU is not saturated, so it is not the bottleneck, which means it is likely not enough Erlang processes have been started.
- Qwertystop 8y agoThe CPU graph peaks near 97% (teal line) at the time when connections-per-second are highest. Are you looking at the red? That's the version without the two patches.
- muststopmyths 8y agooh yeah, you're right. I reversed the two in my head somehow.
- makkesk8 8y agoEven if connections per second can be a magnitude or two lower than requests per second this result is still quite off by today's alternative. 14 core machine comparing .net core with other top webservers: https://www.ageofascent.com/2019/02/04/asp-net-core-saturating-10gbe-at-7-million-requests-per-second/ https://www.ageofascent.com/2019/02/04/asp-net-core-saturati...
- deleted 8y ago[deleted]
- sergiotapia 8y agoThat's really exciting! As someone who dropped out of .NET entirely around the time ASP.Net MVC2 came out, where do you recommend I start looking into aspnet core / .net core? Do you still write core .net in visual studio? or can you use vscode?
- benwilson-512 8y agoA lot of folks are failing to read the article. They're intentionally holding each connection open for 1 whole second. This is a whole different ballgame than benchmarks where each connection is allowed to terminate as rapidly as it can send back a plain text response.
- muststopmyths 8y agoGood point. At first glance, holding the connection open for one second seemed a bit meaningless if they're touting connections/sec. But since they are benchmarking Elixir, there is some amount of overhead involved in that framework's management of connections and requests. If I knew Erlang/Elixir, that would be a fascinating thing to explore. Edit: I'm assuming the saturated CPU comes from Elixir and not the OS. It would be strange for 100k/sec to saturate the TCP stack with 36 cores.
- makkesk8 8y agoTotally missed that. In that case it does make sense.
- 8y ago
- confounded 8y agoIs Elixir/Erlang considered superior to Go for writing high concurrency web servers?
- jadbox 8y agoAs far as from my tests and what I've seen reported online, Go and Rust have a substantial lead (20% ish) over Erlang for high throughput servers. EDIT: I believe this is partially due to Go being a lot more CPU efficient overall than Erlang (see below). So for simple servers, Go and Erlang will match performance, but for slightly more complex web servers that need to crunch some data, Go [and Rust] will outperform the Erlang VM. https://stressgrid.com/blog/benchmarking_go_vs_node_vs_elixir/ https://stressgrid.com/blog/benchmarking_go_vs_node_vs_elixi...
- fermuch 8y agoI would add the detail that for both erlang and elixir, running in one core, multiple cores, or multiple machines is seamless. Clustering is easy.
- StreamBright 8y agoWell one of the fastest HTTP library out there is rapidoid and fasthttp comes very close to it as well as actix-raw, hyper and tokio-minihttp. Erlang and Elixir is lagged behind with a non trivial margin.
- rakoo 8y agoNot na expert in any of the languages by any means, but Go and Erlanger/Elixir focus on different things: - Go wants to be performant at high concurrency scale - Erlang/Elixir wants to keep running at high concurrency scales, whatever the issues are in your application code. Performance comes second. There's no clear cut answer to your question; I guess if you trust yourself to write servers that will hold a large number of connections while doing a lot of processing then Go has an advantage, otherwise you should probably trust the man-centuries behind the BEAM VM and follow the various blog posts/presentations explaining how you can fine-tune your machine to get to super large scales.
- holtalanm 8y agoim a simple man. i see Elixir, i upvote. that being said, this article was pretty informative. The bit about the proposed SO_REUSEPORT socket option was really interesting. Really fun to read about performance bottleneck detection and improvement. edit: wow, downvoting for making a simple joke about liking elixir. Cool.
- mrinterweb 8y agoI've found that humor in comments on HN is usually not well received. Not sure why, just an observation.
- holtalanm 8y agoyeah, i've noticed that, as well. every time i try to make a joke in a comment, it gets downvoted almost immediately. HUMOR? HUMOR HAS NO PLACE HERE! THIS IS A FORUM OF INTELLECTUAL DISCUSSION!
- pmarreck 8y agoIt's ASD. << that was a joke I think that the inclination towards "meaningless" humor makes it too much like Reddit. These folks want SUBSTANCE! (Well, that's why _I_ come here, at least!)
- dang 8y agohttps://news.ycombinator.com/item?id=18817249 https://news.ycombinator.com/item?id=18817249 Maybe we should add something about this to https://news.ycombinator.com/newsfaq.html https://news.ycombinator.com/newsfaq.html.
- thatcat 8y agoSimple jokes from a simple man... I laughed anyway.
- Leace 8y agoejabberd [0], XMPP server is written in erlang and powers chat in some of the biggest MMORPGs [1]. [0]: https://github.com/processone/ejabberd https://github.com/processone/ejabberd [1]: https://xmpp.org/uses/gaming.html https://xmpp.org/uses/gaming.html
- vasilia 8y agoI can handle 120k connections per second with my custom made, highly optimized multiprocess C++ server. But the main problem is business logic. Just make 2 SQL queries to MySQL on each HTTP request and look at how it will degrade.
- jamra 8y agoThat seems like a job for sharded databases and caches.
- vasilia 8y agoOf course, we have both, but 100k is nothing if it's not a CDN server which stores static file in-memory. Moreover, the main metric is latency, not a number of connections. You can scale a number of connections with an L3/L4 load balancer, but not latency.
- repsilat 8y agoThere are simple tricks to make those queries not kill performance. Here is a dumb proof-of-concept I made a few months ago: https://github.com/MatthewSteel/carpool https://github.com/MatthewSteel/carpool The general idea is combining queries from different HTTP requests into a single database query/transaction, amortising the (significant) per-query cost over those requests. For simple use-cases it doesn't add a whole lot of complexity, can reduce both load and latency significantly, and doesn't lose transactional guarantees. Not 100k/sec writes on my laptop, mind you :-).
- chug 8y agoLooks interesting! You mentioned in the docs that it would be simpler once abstractions develop and that made me realize it's similar to facebook/dataloader, just used across requests instead of batching up all of the queries per request. It's also of course a generalized form of it that represents batching a parametrized method more so than just batching retrievals by some kind of unique key. It may be able to serve as something to lift API ideas from though. Like some kind of BatchedTask that has an execute() method that takes an array of args then batches those into an array of array of args for the underlying batched implementation. https://github.com/facebook/dataloader/blob/master/README.md https://github.com/facebook/dataloader/blob/master/README.md
- rargulati 8y agoI'd love to see data on the average on-call incidents for an application written in language X (say Go) vs those written in Elixir. Concretely, its it the case, for an application where Elixir/Erlang/Beam are a great choice, but also, another language would be fine, that the equivalent Elixir application results in less downtime/pages than the alternative. Anything from the perfect app to something with a ton of races/leaks. Is this a fair question (maybe I'm presuming too much of BEAM/supervisor pattern, I zero experience with it)?
- kureikain 8y agoI don't have but I can tell you from my experience with Ruby, Go, Node and Elixir. I have zero on-call for Go. I had very few for Elixir. But the bug were in logic code. Same with Ruby. But it's a disaster with Node. We used TypeScript so it catch lot of type issue. However, the Node runtime is weird. We run into DNS issue(like yo have to bump the libuv thread pool, cache DNS). JSON parsing issue and block the event loop etc...max memory...
- optimusclimb 8y agoThis would be too heavily influenced by confounding factors. For instance: * Are the teams that use certain languages comprised of more experienced people? * How mature is the company and project? I.e., a faster moving startup cutting more corners, where time was decided to be of the essence (rightly or wrongly) will likely produce more on call incidents than a slower, more established company that can takes its time
- rdtsc 8y ago> I'd love to see data on the average on-call incidents Don't have any hard data to compare but having been involved in debugging running Erlang systems. It's very nice having the ability to restart separate supervisors while the rest of the processes handle requests. Being able to do hot code loading to say fix bugs or add extra logging. And my all time favorite -- live tracing after connecting to a VM's remote shell. You can just pick any function, args, and process and say "trace these for a few seconds if a specific condition happens". None of those individually are earth shattering but taken together they are just so pleasant to use. I wouldn't enjoy going back to anything didn't those capabilities. And yes, that restarting of sub-systems (supervision trees) happens automatically as well. There were a number of cases were it turned a potential "wake up 4am and fix this now, cause everything crashed" into a "meh, it's fine until I get to it next week" kind of a problem.
- cutler 8y agoGreat, I can use this for that blogging app I've been meaning to write and sleep at night knowing I won't run out of connections.
- supermatt 8y agoId like to see memory consumption charts for this. It seems you miss this on all your posts. Not a criticism (and thank you for what you have done), its just something I (and others) would like to see, and if you are running the tests its just another metric to log :D Also, any update on your previous article? https://news.ycombinator.com/item?id=19094233 https://news.ycombinator.com/item?id=19094233
- kt315_ 8y agoWe are preparing new benchmark test for major platforms. Among other suggestions it will include memory consumption.
- deleted 8y ago[deleted]
- fabioyy 8y agoOpening a conection and closing after a while is not very good example os "scalability" ... the kernel does the opening part... reply