5 ms·
The article explains that the primary benefit of user-mode green-thread fibers is not switching speed, but that they are much cheaper than OS threads, so you ca
by vii 6y ago
The article explains that the primary benefit of user-mode green-thread fibers is not switching speed, but that they are much cheaper than OS threads, so you can have many more of them. The costs are paid in memory usage and also the operating system also has to do considerable bookkeeping for threads.
However, Netty has offered strong support for callback style IO under the JVM for a long time. This effectively allows the same efficiencies. Of course it is also possible to do without Netty. Therefore the real advantage of Loom user-mode threads and co-routines is syntactical programming convenience. That's the real innovation in Loom!
- free_rms 6y agoI thought the memory cost of threads was more of a JVM thing (default 1mb stack) than an OS thing (pages that aren't mapped/resident don't cost much). What's the cost to the kernel besides a few structs?
- pron 6y agoOnce a page is committed, it cannot be uncommitted until the thread dies, because the OS can't be sure how much of the stack is actually used. It cannot even assume that only addresses above sp are used. Also, the granularity is that of a page, which could be significantly larger than a whole stack of some small, "shallow" thread, and we want lots of small threads.
- marvy 6y agoHow many parked "fibers" do you think a 32-bit JVM might be able to handle once loom is ready for prime time? A thousand? A million? Somewhere in between?
- pron 6y agoWhen somebody shows up to do a 32-bit port we'll think about it :)
- marvy 6y agoI don't see why you need to wait for that. You already know the memory usage for 64-bit, and you implied elsewhere in this thread that memory usage is the key metric that makes these things cheaper than kernel threads. So you should be able to give an order of magnitude guesstimate even if no 32-bit port is in the works yet. Or no?
- pron 6y agoOK, so let's see. Suppose you have 2GB available for threads, and you want a million threads. This leaves 2KB per thread. Can 2KB suffice for a thread? Sure, for some. Could most threads take up less than 2KB? Probably not. Right now the representation of the stack in RAM is not very efficient, but it could be made much more so. If anyone wants to port to 32 bit and that turns out to be a problem, then I guess improving the stack representation would become a higher priority.
- marvy 6y agoThanks! I think that answers my question. Since kernel threads will likely have 64KB stacks or even more, this is still a huge win, even for 32-bit code I guess.
- chris_overseas 6y agoI doubt 32bit is an issue for this and I'd assume well over a million, if Kotlin's coroutines are anything to go by: https://kotlinlang.org/docs/tutorials/coroutines/coroutines-basic-jvm.html#lets-run-a-lot-of-them https://kotlinlang.org/docs/tutorials/coroutines/coroutines-...
- marvy 6y ago32-bit means you have at most 4 gigs of address space to play with; pron implied that lower memory usage is the key savings of fibers vs threads, so I assume a 32-bit JVM will hit a memory limit a lot sooner.
- deleted 6y ago[deleted]
- kgoutham93 6y agoHey Ron, firstly awesome write-up. Can you force yourself to list a few disadvantages of virtual threads over native threads?
- pron 6y agoI wouldn't use virtual threads for long-running, compute-heavy tasks, because even when we expose forced preemption, it might not be as efficient as a kernel preemption. Also, if you run a lot of native (non-Java) code that blocks -- or that upcalls back into Java and then blocks -- then the underlying "carrier" kernel thread will be blocked, and if this happens a lot, it can impact other virtual threads that share the same scheduler. So for compute-heavy jobs and for code that blocks in native or upcalls to Java, the platform threads would be more appropriate.
- jlokier 6y ago> the OS can't be sure how much of the stack is actually used. It cannot even assume that only addresses above sp are used A nitpick, and a possible optimization opportunity. If there are any unmasked signal handlers not using SA_ONSTACK, the OS can reasonably assume the thread doesn't care about memory below sp - redzone (redzone is 128 bytes on AMD64) and therefore there would be no harm in reclaiming it.
- asdfasgasdgasdg 6y agoWell, there is a cap on the number of PIDs you can have going, I think. Also, it is my understanding that thread switching causes TLB flushes, which can be expensive. You can also fit several stackless thread contexts within a single page, depending on the size of the context, whereas you'll certainly never get more than one stackful thread into a single page.
- MaxBarraclough 6y agoI don't think there should be a TLB flush for a switch between threads, as they share a memory space. It's only necessary for switching between processes. https://stackoverflow.com/a/5440165/ https://stackoverflow.com/a/5440165/
- asdfasgasdgasdg 6y agoOops, I think you're right and I was wrong.
- willtim 6y ago> Therefore the real advantage of Loom user-mode threads and co-routines is syntactical programming convenience. That's the real innovation in Loom! It's an interesting alternative to Async-Await / Haskell Monads / F# workflows for sure; and a better fit for Java. But I think some credit is also due for languages like C#, F# and Haskell for providing some healthy competition and prior art in this area.
- pron 6y agoThe relevant prior art is in Scheme (and OCaml, a bit), Erlang and Go, and we do credit them where relevant. Nevertheless, there are a few innovations in Java, both in implementation and design. For example, Java allows you to provide a custom scheduler for virtual threads.
- willtim 6y agoI am very impressed with the recent changes happening to Java, all of which seem very carefully thought out and planned, keep up the good work!
- The_rationalist 6y agoNetty approach should be inferior because of the followings: It doesn't have access to the JVM profiling and semantics It doesn't yet use io uring unlike loom (but there is an active gsoc that should fill this gap) It doesn't use restartable sequences unlike loom. I expect Netty to support Loom instead of their current mechanism, when it becomes available. The thing I wait the most is keeping the standard jdbc API but making it seamlessly truly asynchronous/non socket blocking, that would actually allow Java frameworks to win the TechEmpower benchmarck