13 ms·
Achieving 5M persistent connections with Project Loom virtual threads
- wiseowise 4y agoAnd how is that any different from Kotlin coroutines if you still need to call Thread.startVirtualThread?
- ferdowsi 4y agoKotlin coroutines are colored and infect your whole codebase. Virtual threads do not.
- wiseowise 4y agoYou can mark everything suspend and there's no difference.
- pron 4y ago1. These are actual threads from the Java runtime's perspective. You can step through them and profile them with existing debuggers and profilers. They maintain stacktraces and ThreadLocals just like platform threads. 2. There is no need for a split world of APIs, some designed for threads and others for coroutines (so-called "function colouring"). Existing APIs, third-party libraries, and programs — even those dating back to Java 1.0 (just as this experiment does with Java 1.0's java.net.ServerSocket) — just work on millions of virtual threads. Normally, you wouldn't even call Thread.startVirtualThread(), but just replace your platform-thread-pool-based ExecutorService with an ExecutorService that spawns a new virtual thread for each task (Executors.newVirtualThreadPerTaskExecutor()). For more details, see the JEP: https://openjdk.java.net/jeps/425 https://openjdk.java.net/jeps/425
- pjmlp 4y agoNative VM support instead an additional library faking it, and filling .class files with needless boilerplate.
- invalidname 4y agoThis is pretty fantastic! I'm very excited about the possibilities of Loom. Would love to have a more realistic sample with Spring Boot that would demonstrate the real world scale. I saw a few but nothing remotely as ambitious as that.
- isbvhodnvemrwvn 4y agoSpring Boot overhead would likely make that infeasible.
- invalidname 4y agoI'm not saying 5M. I just want to see to what scale it would get without threading issues. Spring Boot isn't THAT heavy.
- RhodesianHunter 4y agoSpring boot overhead is largely in startup time. It really doesn't have much overhead there after. It's largely a collection of the same libraries you would use anyways glued together with a custom di system.
- nelsonic 4y agoReminds of https://phoenixframework.org/blog/the-road-to-2-million-websocket-connections https://phoenixframework.org/blog/the-road-to-2-million-webs... Would love to see this extended to more Languages/Frameworks.
- bkolobara 4y agoWith lunatic [0] we are trying to bring this to all languages that compile to WebAssembly. A few days ago I wrote about our journey of bringing it to Rust: https://lunatic.solutions/blog/writing-rust-the-elixir-way-1.5-years-later/ https://lunatic.solutions/blog/writing-rust-the-elixir-way-1... [0]: https://github.com/lunatic-solutions/lunatic https://github.com/lunatic-solutions/lunatic
- mike_hearn 4y agoIn theory once Graal adds support for it, any Graal/Truffle-compatible language can benefit. IMHO it's only JVM+Graal that can bring this to other languages. Loom relies very heavily on some fairly unique aspects of the Java ecosystem (Go has these things too though). One is that lots of important bits of code are implemented in pure Java, like the IO and SSL stacks. Most languages rely heavily on FFI to C libraries. That's especially true of dynamic scripting languages but is also true of things like Rust. The Java world has more of a culture of writing their own implementations of things. For the Loom approach to work you need: a. Very tight and difficult integration between the compiler, threading subsystem and garbage collector. b. The compiler/runtime to control all code being used. The moment you cross the FFI into code generated by another compiler (i.e. a native library) you have to pin the thread and the scalability degrades or is lost completely. But! Graal has a trick up its sleeve. It can JIT compile lots of languages, and those languages can call into each other without a classical FFI. Instead the compiler sees both call site and destination site, and can inline them together to optimize as one. Moreover those languages include binary languages like LLVM bitcode and WASM. In turn that means that e.g. Python calling into a C extension can still work, because the C extension will be compiled to LLVM bitcode and then the JVM will take over from there. So there's one compiler for the entire process, even when mixing code from multiple languages. That's what Loom needs. At least in theory. Perhaps pron will contradict me here because I have a feeling Loom also needs the invariant that there are no pointers into the stack. True for most languages but not once C gets involved. I don't know to what extent you could "fix" C programs at the compiler level to respect that invariant, even if you have LLVM bitcode. But at least the one-compiler aspect is not getting in the way.
- notorandit 4y agoWith a maximum of 64k TCP connections per single server IP, you need 77 different IP on the server side. This is a fact.
- imperio59 4y agoPretty sure you can bump that up in the kernel to hold more active connections per server that 64k...
- jauer 4y agoHow do you figure? Clients can connect to the server on the same server port, so connection limit is more like 64k*2 for every Client IP-Server IP pair.
- akvadrako 4y agoActually every client IP+port / server IP+port pair. Linux uses 60999 − 32768 for ephemeral ports so can support 28e3^2 = 784 million connections per IP pair.
- mypalmike 4y agoExcept your service is almost certainly listening on one non-ephemeral port. But having "only" tens of thousands of connections per client is rarely a problem in practice, apart from some load testing scenarios (such as the experiment here, where they opened a number of ports so they could test a large number of connections with a single client machine).
- charcircuit 4y ago1 IP can correspond to multiple different clients.
- deleted 4y ago[deleted]
- 4y ago
- deepsun 4y agoHow does that compare to Kotlin suspend functions?
- torginus 4y agoWhile I can't answer the question directly there is an article about C#-s async/await vs Go's goroutines, which compare the two approaches, and while some of the stuff is probably stack-specific, a lot of it is probably intrinsic to the approach: - Green threads scale somewhat better, but both scale ridiculously well, meaning probably you won't run into scaling issues. - async/await generators use way less memory than a dedicated green thread, this affects both memory consumption and startup time, since the process has to run around asking the OS for more memory - green threads are faster to execute Here's the link: https://alexyakunin.medium.com/go-vs-c-part-1-goroutines-vs-async-await-ac909c651c11 https://alexyakunin.medium.com/go-vs-c-part-1-goroutines-vs-...
- jillesvangurp 4y agoLoom will make a great backend for kotlin's co-routines. Roman Elizarov (kotlin language lead & person who is behind Kotlin's co-routine framework) has already confirmed that will happen and it makes a lot of sense. For those who don't understand this, Kotlin's co-routine framework is designed to be language neutral and already works on top the major platforms that have kotlin compilers (native, javascript, jvm, and soon wasm). So, it doesn't really compete with the "native" way of doing concurrent, aynchronous, or parallel computing on any of those platforms but simply abstracts the underlying functionality. It's actually a multi platform library that implements all the platform specific aspects in the platform appropriate way. It's also very easy to adapt existing frameworks in this space via Kotlin extension functions and the JVM implementation actually ships out of the box with such functions for most common solutions on the JVM for this (Java's threads, futures, threadpools, etc., Spring Flux, RxJava, Vert.x, etc.). Loom will be just another solution in this long list. If you use Spring Boot with Kotlin for example, rather than dealing with Spring's Flux, you simply define your asynchronous resources as suspend functions. Spring does the rest. With Kotlin-js in a browser you can call Promise.toCoroutine() ans async { ... }.asPromise(). That makes it really easy to write asynchronous event handling in a web application for example or work with javascript APIs that expect promises from Kotlin. And if you use web-compose, fritz2, or even react with kotlin-js, anything asynchronous, you'd likely be dealing with via some kind of co-routine and suspend functions. Once Loom ships, it basically will enable some nice, low level optimization to happen in the JVM implementation for co-routines and there will likely be some new extension functions to adapt the various new Java APIs for this. Not a big deal but it will probably be nice for situations with extremely large amounts of co-routines and IO. Not that it's particularly struggling there of course but all little bits help. It's not likely to require any code updates either. When the time comes, simply update your jvm and co-routine library and you should be good to go.
- KingOfCoders 4y agoSomething to learn for everybody, the article is mainly about Linux tuning.
- jeroenhd 4y agoThe Linux tuning part seems to have been inspired by these blog posts from 14 years ago: https://www.metabrew.com/article/a-million-user-comet-application-with-mochiweb-part-1 https://www.metabrew.com/article/a-million-user-comet-applic... It's almost a little disappointing that beefy modern servers only manage a x5 scale improvement, though that could be due to the differences in runtime behaviour between Erlang and the JVM.
- toast0 4y agoI mean... is 5M very impressive? Not really. Does it show that Project Loom meets the goal of being able to do large client count thread per server workloads? I think so. Does the name remind me of a best selling point and click adventure game? Definitely yes.
- wiradikusuma 4y agoThe experiment is about Java app, but the tweaks are at the O/S level. Does it mean any app (Java/not, Loom/not) can achieve target given correct tweak? Also, why are these not default for the O/S? What are we compromising by setting those values?
- jiggawatts 4y agoThere's always trade-offs. It would be very rare for any server to reach even 100K concurrent connections, let alone 5M. Optimising for that would be optimising for the 0.000001% case at the expense of the common case. Some back of the envelope maths: https://www.wolframalpha.com/input?i=100+Gbps+%2F+5+million https://www.wolframalpha.com/input?i=100+Gbps+%2F+5+million If the server had a 100 Gbps Ethernet NIC, this would leave just 20 kbps for each TCP connection. I could imagine some IoT scenarios where this might be a useful thing, but outside of that? I doubt there's anyone that wants 20 kbps throughput in this day and age... It's a good stress test however to squeeze out inefficiencies, super-linear scaling issues, etc...
- Koffiepoeder 4y agoOpen, idle websockets can be a use case for a large amount of tcp connections with a small data footprint.
- jeffbee 4y agoAlso IMAP has this unfortunate property.
- jeroenhd 4y ago20kbps should be sufficient for things like chat apps if you have the CPU power to actually process chat messages like that. Modern apps also require attachments and those will require more bandwidth, but for the core messaging infrastructure without backfilling a message history I think 20kbps should be sufficient. Chat apps are bursty, after all, leaving you with more than just the average connection speed in practice.
- torginus 4y agoWhile impressive, I don't really see it as something practical - I think scaling across processes/VMs is a much more realistic approach.
- cheradenine_uk 4y agoI think a lot of people are missing the point. Go look at the sourcecode. Look at how simple it is - anyone who has created a thread with java knows what's happening. With only minor tweaks, this means your pre-existing code can take advantage of this with, basically, no effort. And it retains all the debuggability of traditional java thread (I.e: a stack trace that makes sense!) If you've spent any time at all dealing with the horrors of c# async/await (Why am I here? Oh, no idea) and it's doubling of your APIs to support function colouring - or, you've fought with the complexities of reactive solutions in the Java space -- often, frankly, in the name of "scalability" that will never be practically required -- this is a big deal. You no longer have to worry about any of that.
- bullen 4y agoAgreed it's simpler, but using NIO with one OS thread per core also has it's benefits. The context switch (how ever small) will cause latency when this solution is at saturation. I think they should write four tests: fiber, NIO and each with userspace networking (no kernel copying network memory) and compare them. Why Oracle is stalling removing the kernel for Java networking is surprising to me, they allready have a VM.
- pron 4y agohttps://github.com/ebarlas/project-loom-comparison https://github.com/ebarlas/project-loom-comparison
- vlovich123 4y agoShouldn’t you be able to send authorization and authentication requests in parallel in the async and virtual threads cases?
- threeseed 4y agoIt is just an example so they could do anything. But in the real world it is common to need information from the authorization stage to use in the authentication stage. For example you may have a user login with an email address/password which you then pass to an LDAP server in order to get a userId. This userId is then used in a database to determine with objects/groups they have access to.
- the8472 4y agonet.netfilter.nf_conntrack_buckets = 1966050 net.netfilter.nf_conntrack_max = 7864200 or avoid conntrack entirely
- LinuxBender 4y agoFor completeness sake I would add that one must also set options nf_conntrack expect_hashsize=X hashsize=X in /etc/modules.d/nf_conntrack.conf, X being 1/4 the size of conntrack_max
- pron 4y agoFor more information about virtual threads see https://openjdk.java.net/jeps/425 https://openjdk.java.net/jeps/425 (planned to preview in JDK 19, out this September). What's remarkable about this experiment is that it uses simple 26-year-old (Java 1.0) networking APIs.
- newskfm 4y ago
- zinxq 4y agoLoom sets out to give you a sane programming paradigm similar to what threads do (i.e. as opposed to programming asynchronous I/O in Java with some type of callback) without the overhead of Operating System threads. That's a very cool and a noble pursuit. But the title of this article might as well have been "5M persistent connections with Linux" because that's where the magic 5M connections happen. I could also attempt 5M connections at the Java level using Netty and asynchronous IO - no threads or Loom. Again, it'd take more Linux configuration than anything else. If that configuration did happen though now you can also do it in C# async/await, javascript, I'm sure Erlang and anything else that does Asynchronous I/O whether it's masked by something like Loom/Async/Await or not.
- pron 4y agoIt is true that the experiment exercises the OS, but that's only part of the point. The other part is that it uses a simple, blocking, thread-per-request model with Java 1.0 networking APIs. So this is "achieving 5M persistent connections with (essentially) 26-year-old code that's fully debuggable and observable by the platform." This stresses both the OS and the Java runtime. So while you could achieve 5M in other ways, those ways would not only be more complex, but also not really observable/debuggable by Java platform tools.
- cheradenine_uk 4y agoThis. Writing the sort of applications that I get involved with, it's frequently the case whilst it's true that 1 OS thread/java thread was a theoretical scalability limitation - in practice we were never likely to hit it (and there was always the 'get a bigger computer'). But: the complexity mavens inside our company and projects we rely upon get bitten by an obsessive need to chase 'scalability' /at all costs/. Which is fine, but the downside to that is the negative consequences of coloured functions comes into play. We end up suffering having to deal with vert.x or kotlin or whatever flavour-of-the-month solution is that is /inherently/ harder to reason about than a linear piece of code. If you're in a c# project, the you get a library that's async, and boom, game over. If loom gets even within performance shouting distance of those other models, it's ought to kill (for all but the edgiest of edge-cases) reactive programming in the java space dead. You might be able to make a case - obviously depending on your use cases which are not mine - that extracting, say, 50% more scalability is worth the downsides. If that number is, say, 5%, then for the vast majority of projects the answer is going to be 'no'. I say 'ought to', as I fear the adage that "developers love complexity the way moths love flames - and often with the same results". I see both engineers and projects (Hibernate and keycloak, IIRC) have a great deal of themselves invested in their Rx position, and I already sense that they're not going to give it up without a fight. So: the headline number is less important than "for virtually everyone you will no longer have to trade simplicity for scalability". I can't wait!
- alberth 4y agoIs this a test of just having 5M people knock on your door? Or is this a test where something actually happens (data exchanges) with each connection? I ask because those are two totally different workloads and typically where in the later test Erlang shines.
- bufferoverflow 4y agoIt's an echo server. The client sends the data, the server responds with the same data.
- 10000truths 4y agoA bit of a digression, but I’d love to see how much further one could go with a memory-optimized userland TCP stack, and storing the send and receive buffers on disk. A TCP connection state machine consists of a few variables to keep track of sequence numbers and congestion control parameters (no more than 100-200 bytes total), plus the space for send/receive buffers. A 4 TB SSD would fit ~125 million 16-KB buffer pairs, and 125 million 256-byte structs would take up only 32 GB of memory. In theory, handling 100 million simultaneous connections on a single machine is totally doable. Of course, the per-connection throughput would be complete doodoo even with the best NICs, but it would still be a monumental yet achievable milestone.
- mike_hearn 4y agoPresumably at 100M simultaneous connections the machine CPU would be saturated with setting up and closing them, without getting much actual work done. TCP connections seem too fragile to make it worth trying to keep them open for really long periods. It's interesting to think about though, I agree. What are the next scaling bottlenecks now (for JVM compatible languages) threading is nearly solved? There are some obvious ones. Others in the thread have pointed out network bandwidth. Some use cases don't need much bandwidth but do need intense routability of data between connections, like chat apps, and it seems ideal for those. Still, you're going to face other problems: 1. If that process is restarted for any reason that's a lot of clients that get disrupted. JVMs are quite good at hot-reloading code on the fly, so it's not inherently the case that this is problematic because you could make restarts very rare. But it's still a problem. 2. Your CPU may be sufficient for the steady state but on restart the clients will all try to reconnect at once. Adding jitter doesn't really solve the issue, as users will still have to wait. Handling 5M connections is great unless it takes a long time to reach that level of connectivity and you are depending on it. 3. TCP is rarely used alone now, it usually comes with SSL. Doing SSL handshakes is more expensive than setting up a TCP connection (probably!). Do you need to use something like QUIC instead? Or can you offload that to the NIC making this a non-issue? I don't know. BTW the Java SSL stack is written in Java itself so it's fully Loom compatible.
- charcircuit 4y ago
- Nullabillity 4y agoLoom is missing the point. Time has shown that bare threads are not a viable high-level API for managing concurrency. As it turns out, we humans don't think in terms of locks and condvars but "to do X, I first need to know Y". That maps perfectly onto futures(/promises). And once you have those, you don't need all the extra complexity and hacks that green threads (/"colourless async") bring in. I'd take a system that combined the API of futures with the performance of OS threads over the opposite combination, any day of the week. But as it turns out, we don't have to choose. We can have the performance of futures with the API of futures. Or we can waste person-years chasing mirages, I guess. I just hope I won't get stuck having to use the end product of this.
- rvcdbn 4y agoMaybe threads don’t work for your thinking style but your claim that this is generally true is baseless and pretty well refuted by languages like Go or Erlang that feature stackfull threads/processes as a critical part of their best-in-class concurrency stories.
- Nullabillity 4y agoErlang sidesteps the problem by avoiding mutable shared state, in this context they're threads/processes in name only. Go is just yet another implementation of green threads that is slightly less broken than prior implementations, because it had the benefit of being implemented on day 1 (so the whole ecosystem is green thread-aware). It's certainly nowhere near "best-in-class".
- chrisseaton 4y ago> Erlang sidesteps the problem by avoiding mutable shared state Erlang is maximal shared mutable state! Processes are mutable state and they’re shared between other processes.
- toast0 4y agoShared mutable state is hard to work with, but Java threads and Java promises both give you access to it. In either case, you'd need discipline to avoid patterns which reduce concurrency. From the article, it seems that Loom (in preview) enables the threaded model for Java to scale. IMHO, this is great because you can write simple straightforward code in a threaded model. You can certainly write complex code in a threaded model too. Maybe there's an argument that promises can be simple and straightforward too, but my experience with them hasn't been very straightforward.
- christophilus 4y agoLoom looks like it’s nicely solved the function coloring problem. This plus Graal makes me excited to pick up Clojure again.
- imranhou 4y agoIt looks more closer to go routines, which to me begs the question - where are the channels that I could use to communicate between these virtual threads?
- sdfgdfgbsdfg 4y agoIn a library. Loom is more about adapting the JVM itself for continuations and virtual threads than adding to userspace.
- deleted 4y ago[deleted]
- adra 4y agoGo's channels are simplistically a mutex in front of a queue. Java has many existing objects that can do the same, it's just that's not idiomatic best choice to do the same. Since green threads should wake up from Object.notify(), any threads blocking on the monitor should wake/consume. I'm curious how scalable/performance a green thread ConcurrentDequeue would stand up to go's channel.
- Matthias247 4y agoYou are right. But Go Channels come also with the superpower of „select“, which allows to wait for multiple objects to become ready and atomic execution of actions. I don’t think this part can be retrofitted on top of simple BlockingQueues.
- sdfgdfgbsdfg 4y agopron talks about this on https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part2.html#channels https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part2....
- TYMorningCoffee 4y agoI was only able to get to 840,000 open connections with my experiment. My machine only has 8GB of memory. https://josephmate.github.io/2022-04-14-max-connections/ https://josephmate.github.io/2022-04-14-max-connections/ Is there anyway for the TCP connections share memory in kernel space? My experiment only uses two 8 byte buffers in userspace.
- mh- 4y agono*, and as you've discovered, the skbufs allocated by the kernel will often be the limiting factor for a highly concurrent socket server on linux. * I don't know if someone has created some experimental implementation somewhere. It would require a significant overhaul of the TCP implementation in the kernel. edit: check out this sibling thread about userland TCP. I think this is a more interesting/likely direction to explore in. https://news.ycombinator.com/item?id=31215569 https://news.ycombinator.com/item?id=31215569
- toast0 4y agoDoes Linux actually allocate buffers for each socket or does it just link to sk_buff's (which I understand are similar to FreeBSD's mbuf's) and then limit how much storage can be linked? FreeBSD has a limit on the total ram used for mbufs as well, not sure about Linux. Otoh, FreeBSD's maximum FD limit is set as a factor of total memory pages (edit: looked it up, it's in sys/kern/subr_param.c, the limit is one FD per four pages, unless you edit kernel source) and you've got 2M pages with 8GB ram, so you would be limited to 512k FDs total, and if you're running the client on the same machine as server, that's 256k connections. But 8G is not much for a server, and some phones have more than that... so it's not super limiting. When you're really not doing much with the connections, userland tcp as suggest in a sibling, could help you squeeze in more connections, but if you're going to actually do work, you probably need more ram. Btw, as a former WhatsApp server engineer, WhatsApp listens on three ports; 80, 443, and 5222. Not that that makes a significant difference in the content.
- Matthias247 4y agoI think the socket buffers (sk_buff) are actually shared. They are all packet sized, and whatever socket needs to transmit some data or receives it gets the buffers attached. So my assumption is that the amount of required socket buffers scales more with the amount of data transmission than with the number of sockets. But independent of socket buffers, the kernel obviously needs to allocate other state per socket, which tracks the state of the TCP connection.
- Andrew_nenakhov 4y agoSounds like a job for Erlang.
- speed_spread 4y agoSounds like Erlang's out of a job.
- Andrew_nenakhov 4y agoNo.
- sgtnoodle 4y agoI'm not a java programmer. I tried clicking 3 layers deep of links, but still have no idea what virtual threads are in this context. Is it a userspace thread implementation? I've used explicit context switching syscalls to "mock out" embedded real time OS task switching APIs. It's pretty fun and useful. The context switching itself may not be any faster than if the kernel does it, but the fact that it's synchronous to your program flow means that you don't have to spend any overhead synchronizing to mutexes, queues, etc. (You still have them, they just don't have to be thread safe.)
- grishka 4y ago> Is it a userspace thread implementation? Yes.
- metabrew 4y agoAPI for the server example looks... actually good, wow. Nice job! Also tickled to see my erlang 1M comet blog post referenced. A lifetime ago now, pre-websockets.
- midislack 4y agoI see a lot of these making the FP of HN. But it's very difficult to be impressed, or unimpressed because it's all about hardware. How much hardware is everybody throwing at all of this? 5M persistent connections on a Pi with mere GigE? Pretty frickin' amazing. 5M persistent connections on a Threadripper with 128 cores and a dozen trunked 4 port 10GE NICs? Yaaaaawwwnnn snooze. We need a standardized computer for benchmarking these types of claims. I propose the RasPi 4 4GB model. Everybody can find one, all the hardware's soldered on so no cheating is really possible, etc. Then we can really shoot for efficiency.
- shadowpho 4y agoRaspberry pi 4 performance changes wildly based on cooling. Bare die vs heatsink vs heatsink + fan will give you wildly different results.
- niederman 4y ago> Everybody can find one LMAO I wish. https://rpilocator.com/?cat=PI4 https://rpilocator.com/?cat=PI4
- kmelva 4y agoCould a 128c Threadripper even do 5M kernel threads?
- jpollock 4y agoThis isn't about the hardware, it's about thread count. There are limits in the linux kernel, and the 5m concurrent connections was chosen to exceed it. From what I remember (my knowledge is ancient though), a Java thread consumes a pid_t in the linux kernel. By default this is limited to 64k. However, this can be increased by setting a flag in the kernel, to a maximum 2^22 or 4m. In order to have more than 4m connections, the existing Java code either needs to be changed to be event driven, or it can't use kernel threads. Event driven code is very different. It's very powerful, but it is very easy to get lost. Think writing Java code that looks like a Makefile with dependencies or "andThen" everywhere, and everyone having to make sure everything is threadsafe. Thread safety is hard for large teams with high qps services - deadlocks can bring down a service. If a developer can write "regular" non-re-entrant Java code and still get the concurrent connections? Win all around.