10 ms·
Thread Pools on the JVM
- cogman10 5y agoLoom can't land fast enough! The current issue the JVM has is that all threads have a corresponding operating system thread. That, unfortunately, is really heavy memory wise and on the OS context switcher. Loom allows java to have threads as light weight as a goroutine. It's going to change the way everything works. You might still have a dedicated CPU bound thread pool (the common fork join pool exists and probably should be used for that). But otherwise, you'll just spin up virtual threads and do away with all the consternation over how to manage thread pools and what a thread pool should be used for.
- bestinterest 5y agoWhats the difference between goroutines and project loom? Is their any?
- cogman10 5y agoTerminology mostly :D I've not looked into the goroutine implementation, so I couldn't tell you how it compares to what I've read loom is doing. Loom is looking to have some extremely compact stacks which means each new "virtual thread" as they are calling them will end up having bytes worth of memory allocated. Another thing coming with loom that go lacks is "structured concurrency". It's the notion that you might have a group of tasks that need to finish before moving on from a method (rather than needing to worry about firing and forgetting causing odd things to happen at odd times).
- jayd16 5y ago>structured concurrency That's good to hear. You see a lot of these Loom discussions talk about implicit and magical asynchronous execution. I was afraid fine grained thread control would be left out. Its super useful if you want to interface with how most GUI frameworks function (ie a Main thread), or important OS threads like a thread with a bound GL context or what have you.
- cogman10 5y agoYeah, while virtual threads are the bread and butter of Loom, they are also adding a lot of QoL things. In particular, the notion of "ScopedVariables" will be a godsend to a lot of concurrent work I do. It's the notion of "I want this bit of context to be carried through from one thread of execution to the next". Beyond that, one thing the loom authors have suggested is that when you want to limit concurrency the better way to do that is using concurrency constructs like semaphores rather than relying on a fixed pool size.
- ccday 5y agoNot sure if it counts as structured concurrency but Go has the feature you describe: https://gobyexample.com/waitgroups https://gobyexample.com/waitgroups
- jayd16 5y agoThe biggest difference is probably that the JVM will support both OS and lightweight threads. That's really useful for certain things talking to the GPU in a single thread context.
- _old_dude_ 5y agoUnlike go routine, Loom virtual threads are not preempted by the scheduler. I believe you may be able to explicitly preempt a virtual thread but the last time i checked it was not part of the public API
- vips7L 5y agoUnless I'm misunderstanding, virtual threads are preemptive: https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part1.html https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part1....
- _old_dude_ 5y agoBy the OS, not by the scheduler see https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part2.html#forced-preemption https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part2....
- vips7L 5y agoWhat about pron's comments here then? https://news.ycombinator.com/item?id=27885569 https://news.ycombinator.com/item?id=27885569 > Second, Loom's virtual threads can also be forcibly preempted by the scheduler at any safepoint to implement time sharing
- _old_dude_ 5y agoFor me, preemption by the Java scheduler is not currently supported but may be added in the future, after all the goroutine were not preempted at the beginning in Go. The whole quote > Second, Loom's virtual threads can also be forcibly preempted by the scheduler at any safepoint to implement time sharing. Currently, this capability isn't exposed because we're yet to find a use-case for it I believe it's a reference to [1] but i may be wrong. [1] https://download.java.net/java/early_access/loom/docs/api/java.base/java/lang/Continuation.html#tryPreempt(java.lang.Thread) https://download.java.net/java/early_access/loom/docs/api/ja...
- ovis 5y agoWhat benefits does loom provide vs using something like cats-effect fibres?
- _old_dude_ 5y agoYou can actually debug the code you write because you get a real stacktrace, not few frames that shows the underlying implementation.
- clhodapp 5y agoAdmittedly, loom will do much better but cats-effect does try its best within the limitations of the current JVM: https://typelevel.org/cats-effect/docs/2.x/guides/tracing https://typelevel.org/cats-effect/docs/2.x/guides/tracing
- Nullabillity 5y agoOn the other hand, you'll spend a lot more time debugging Loom code, because it reuses the same broken-by-design thread API.
- elygre 5y agoWhat is broken-by-design about the api?
- Nullabillity 5y agoFundamentally, an async API is either data-oriented (Futures/Promises: tell me what data this task produced) or job-oriented (Threads: tell me when this task is done). You can think of it like functions vs subroutines. Since you typically care about the data produced by the task, threads require you to sort out your own backchannel for communicating this data back (such as: a channel, a mutexed variable, or something else). Unscientifically speaking, getting this backchannel wrong is the source of ~99% of multithreading bugs, and they are a huge pain to fix. You can implement futures on top of threads by using a thread + oneshot channel, but that requires that you know about it, and keep them coupled. The point of futures is that this becomes the default correct-by-default API, unless someone goes out of their way to do it some other way. On the other hand, implementing threads on top of futures is trivial: just return an empty token value. There are also some performance implications: depending on your runtime it might be able to detect that future A is only used by future B, and fuse them into one scheduling unit. This becomes harder when the channels are decoupled from the scheduling.
- christkv 5y agoAre we coming full circle going back a variant of the original Java green threads?
- hashmash 5y agoNot quite. The original green threads were seen as more of a hack until Solaris supported true threads. Green threads could only support one CPU core, and so without a major redesign, it was a dead end.
- AtlasBarfed 5y agoBasically yes. Longer answer: devs back in the day couldn't really grok the difference between green and real threads. Java made its bones as an enterprise language, which can have smart programmers, but they will conversely not be closer-to-metal knowledgewise. Too many devs back in the day expected a java thread to be a real thread, so java re-engineered to accomodate this. I think the JDK/JVM teams also viewed it as a maturation of the JVM to be directly using OS resources so closely across platforms, rather than "hacking" it with green threads. These days, our high performance fanciness means the devs are demanding green thread analogues, and go/elixir/others are seemingly superior because of those. So to remain competitive in the marketplace, Java now needs threads that aren't threads even though Java used to have threads that weren't threads.
- cogman10 5y agoYes and no. The new Loom threads will be much lighter weight than the original Java green threads. Further, the entire IO infrastructure of the JVM is being reworked for Loom to make sure the OS doesn't block the VM's thread. What's more, Loom does M:N threading. Same concept, very different implementation.
- iamcreasy 5y agoSo, with Loom now we can tell exactly in which order theses threads were executed as it's not up to OS to decide thread execution order anymore?
- Spivak 5y agoYou are ignoring the downside to green threads which is that it’s cooperative. If the thread doesn’t yield control back to the event loop then the real OS thread backing the loop is now stuck. Which leads to dirty things like inserting sleep 0 at the top of loops and dealing with really unbalanced scheduling of threads don’t hit yields often enough. Plus with loom it might not be obvious that some function is a yield since it’s meant to be transparent so if you grab a lock and yield you make everyone wait until your scheduled again. Green threads are great! I love them and they’re the only real solutions to really concurrent IO heavy workloads but it’s not a panacea and trades one kind of discipline for another.
- sudhirj 5y agoSleep 0 sounds like quite a hack, Go has the neater https://pkg.go.dev/runtime#Gosched https://pkg.go.dev/runtime#Gosched instead, and I assume there will be a Java equivalent as well. And if most stdlib methods and all blocking methods call it, it's going to be pretty difficult to hang a green thread.
- WatchDog 5y agoFWIW, Java has had `Thread#yield()`[0] since inception. [0]: https://docs.oracle.com/javase/7/docs/api/java/lang/Thread.html#yield https://docs.oracle.com/javase/7/docs/api/java/lang/Thread.h...()
- kaba0 5y agoSince there is a runtime that knows everything about the state of the thread, my understanding is that there is no need for explicit yields. Everything will turn automagically into non-blocking (except for FFI)
- brokencode 5y agoI was under the impression that Loom was implementing preemptable lightweight threads. Is that not the case?
- clhodapp 5y ago
- jeffbee 5y agoAre you quite certain that a (linux, nptl) thread costs more memory than a goroutine? You've implied that but it's not obviously true.
- dragontamer 5y agoWouldn't any linux/nptl thread require at at least the register-state of the entire x86 (or ARM) CPU? I don't think goroutines would need such information. A goroutine knows that "int foobar;" is currently being stored in "rbx", and that "int foobar" is currently saved on the stack. Therefore, rbx doesn't need to be saved. ------ Linux/NPTL threads don't know when they are interrupted. So all register state (including AVX512 state if those are being used) needs to be saved. AVX512 x 32 is 2kB alone. Even if AVX512 isn't being used by a thread (Linux detects all AVX512 registers to be all-zero), RAX through R15 is 128-bytes, plus SSE-registers (another 128-bytes) or ~256 bytes of space that the goroutines don't need. Plus whatever other process-specific information needs to be saved off (CPU time and other such process / thread details that Linux needs to decide which threads to process next)
- jeffbee 5y agoI don't think the question is dominated by machine state, I think it would be more of a question of stack size. They are demand-paged and 4k by default for native threads, 2k by default for goroutines but stored on a GC'd heap that defaults to 100% overhead, so it sounds like a wash to me.
- dragontamer 5y agoHmmm. It seems like you're taking this from a perspective of "Pthreads in C++ vs Coroutines in Go", which is correct in some respects, but different from how I was taking the discussion. I guess I was taking it from a perspective of "pthreads in C++ vs Go-like coroutines reimplemented in C++", which would be pthreads vs C++20 coroutines. (Or really: it seems like this "Loom" discussion is more of a Java thing but probably a close analog to the PThreads in C++ vs C++20 Coroutines) I agree with you that that the garbage collector overhead is a big deal in practice. But its an aspect of the discussion I was purposefully avoiding. But I'm also not the person you responded to.
- cbsmith 5y ago> That, unfortunately, is really heavy memory wise and on the OS context switcher. So, there was a time where a broad statement like that was pretty solid. These days, I don't think so. The default stack size (on 64-bit Linux) is 1MB, and you can manipulate that to be smaller if you want. That's also the virtual memory. The actually memory usage depends on your application. There was a time where 1MB was a lot of memory, but these days, for a lot of contexts, it's kind of peanuts unless you have literally millions of threads (and even then...). Yes, you can be more memory efficient, but it wouldn't necessarily help that much. Similarly, at least in the case of blocking IO (which is normally why you'd have so many threads), the overhead on the OS context switcher isn't necessarily that significant, as most threads will be blocked at any given time, and you're already going to have a context switch from the kernel to userspace. Depending on circumstance, using polling IO models can lead to more context switching, not less. There's certainly circumstances where threads significantly impede your application's efficiency, but if you are really in that situation you likely already know it. In the broad set of use cases though, switching from a thread-based concurrency model to something else isn't going to be the big win people think it will be.
- user5994461 5y ago>>> The default stack size (on 64-bit Linux) is 1MB The default thread stack size is 8 or 10 MB on most Linux. The exception is Alpine that's below 1 MB.
- ori_b 5y agoThe default reserved size is 8mb. The allocated size starts at a page (usually 4k), and grows in page sized increments as you use it.
- cbsmith 5y agoTo clarify, the 1MB is the default stack size for threads with the JVM on 64-bit Linux. Search for "-Xss": https://docs.oracle.com/en/java/javase/16/docs/specs/man/java.html https://docs.oracle.com/en/java/javase/16/docs/specs/man/jav...
- kllrnohj 5y ago
- lmilcin 5y agoI have discovered ReactiveX for Java and Reactor in particular. I am working with Kafka and MongoDB and it is normal for my app to have a million in flight transactions at various stages of completion. In the past it required a lot of planning (and a lot of code) but Reactor let's me build these processes as pipelines with whatever concurrency or scheduler I desire, at any stage of the processing. We are even doing tricks like merging unrelated queries to MongoDB so that sometimes thousands of same queries are executed together (one query with huge in() or one bulk write rather than separate ones). This is improving our throughputs by orders of magnitude while the pipeline pulls millions of documents per second from the database. I just don't see how Loom helps. Loom could help if you had blocking APIs to start, but you get much better results if you just resolve to use async, non-blocking wrapped in ReactiveX.
- geodel 5y agoLoom will help folks who prefer writing straightforward Java code instead of some random reactive library with obscure exception handling and poor to impossible debuggability. Now I get it is hard for many folks to understand that part. Just like at my workplace people think it is impossible to write micro service without SpringBoot. > Loom could help if you had blocking APIs to start, but you get much better results if you just resolve to use async, non-blocking wrapped in ReactiveX. There might be billions of lines of legacy code which would adapt to Loom with minimal changes but impossible to turn in ReactiveX etc without enormous investment and risk. Your ideas are rather simplistic for real world.
- deleted 5y ago[deleted]
- dikei 5y agoYup, Loom will simplify a lot the Producer-Consumer pattern on I/O operation. With virtual threads, it's basically free to block on consumer threads, so you would need only 1 bounded pool for the consumers. Currently for efficiency, you would need at least 2 pools: 1 small bounded pool for dequeuing the requests and create the IO operation, and 1 unbounded pool for actually executing the IO operation.
- 0xffff2 5y agoThis seems like good advice in general. Is any of it really specific to the JVM? If I was doing thread pooling with CPU and IO bound tasks, I would approach threading in a similar way in C++.
- cogman10 5y agoIt'll depend on if your language has either coroutines or lightweight threads. Threadpooling only matters if you have neither of those things. Otherwise, you should be using one or the other over a thread pool. You might still spin up a threadpool for CPU bound operations, but you wouldn't have one dedicated to IO. As of C++ 20, there are coroutines which you should be looking at (IMO). https://en.cppreference.com/w/cpp/language/coroutines https://en.cppreference.com/w/cpp/language/coroutines
- dragontamer 5y agoThreadpools are probably better on CPU-bound bound (or CPU-ish bound tasks: like RAM-bound) without any I/O. Coroutines / Goroutines and the like are probably better on I/O bound tasks where the CPU-effort in task-switching is significant. -------- For example: Matrix Multiplication is better with a Threadpool. Handling 1000 simultaneous connections when you get Slashdotted (or "Hacker News hug of death") is better solved with coroutines.
- cogman10 5y agoI agree. Coroutines MIGHT be more efficient if what you end up building is a statemachine anyways (as that's what most of those coroutines are doing with the compiler). Otherwise, if it's just pure parallel CPU/memory burning with little state transitions/dependence then a dedicated CPU pool fixed to roughly the number of CPU cores on the box will be the most efficient. Heck, it can often even yield benefits to "pin" certain tasks to a thread to keep the CPU cache filled with relent data. For example, 4 threads handling the 4 quadrants of the matrix rather than having the next available thread picking up the next task.
- 5y ago
- jfoutz 5y agoI'm wary of unbounded thread pools. Production has a funny way of showing that threads always consume resources. A fun example is file descriptors. An unexpected database reboot is often a short outage, but it's crazy how quickly unbounded thread pools can amplify errors and delay recovery. Anyway, they have their place, but if you've got a fancy chain of micro services calling out to wherever, think hard before putting those calls in an unbounded thread pool.
- sk5t 5y agoAnd you should be wary! Prefer instead a bounded thread pool with a bounded queue of tasks waiting for service, and also decide explicitly what should happen when the queue fills up or wait times become too high (whatever "too high" means for the application).
- jeffbee 5y agoUnbounded thread pools are bad, bounded thread pool executors with unbounded work queues are bad, and bounded thread pools with bounded queues, FIFO policies, and silent drops are also bad. There are many bad ways to do this.
- dimitrov 5y ago> and bounded thread pools with bounded queues, FIFO policies, and silent drops are also bad. Care to elaborate please? Seems like the author is recommending unbounded thread pools with bounded queues for blocking IO. Isn't that pretty similar?
- jfoutz 5y agoI can't speak for the parent, some things that stand out to me 1. k8s and bare metal, when you make a bunch of threads things get slower. with the FIFO case, you can have pending requests in the queue that don't get their connection canceled event, and the same user puts another request in the queue. 2. Silently dropping is bad, you want an alert - really you want an alert when you get close, so you can add more capacity 3. bounded queue with unbounded threads is really just an unbounded queue - a short line with a mob pushing to get in line Then, you know, memory on k8s, pod gets OOM killed. that sucks cause you have to reschedule and restart. all the pending requests are dropped. It's very easy to make something that works, but is actually quite detrimental when things are on fire. little extra gasoline helps get over the hills, but when things are on fire, gasoline makes a bigger fire.
- jackcviers3 5y agoAuthor mentions scala. Both ZIO[1] and Cats-Effect[2] provide fibers (coroutines) over these specific threadpool designs today, without the need for Project Loom, and give the user the capability of selecting the pool type to use without explicit reference. They are unusable from Java, sadly, as the schedulers and ExecutionContexts and runtime are implicitly provided in sealed companion objects and are therefore private and inaccessible to Java code, even when compiling with ScalaThenJava. Basically, you cannot run an IO from Java code. You can expose a method on the scala side to enter the IO world that will take your arguments and run them in the IO environment, returning a result to you, or notifying some Java class using Observer/Observable. This can, of course take Java lambdas and datatypes, thus keeping your business code in Java should you so desire. It's clunky, though, and I wish Java had easy IO primitives like Scala. 1. https://github.com/zio/zio https://github.com/zio/zio 2. https://typelevel.org/cats-effect/versions https://typelevel.org/cats-effect/versions
- rzzzt 5y agoQuasar has similar functionality: https://docs.paralleluniverse.co/quasar/ https://docs.paralleluniverse.co/quasar/
- cogman10 5y agoFun fact, one of the primary loom devs wrote quasar.
- AzzieElbab 5y agoThat gist is from D.J. Spiewak - one of the authors of cats effect :)
- charleslmunger 5y agoAnother tip - If you have a dynamically-sized thread pool, make it use a minimum of two threads. Otherwise developers will get used to guaranteed serialization of tasks, and you'll never be able to change it.
- bobbylarrybobby 5y agohttps://www.hyrumslaw.com https://www.hyrumslaw.com
- hellectronic 5y agonice!
- u678u 5y agoWith Python at first I was scared of GIL being single threaded, now I'm used to it and it works great. Thousands of threads used to be normal for my old Java projects but seems crazy to me now.
- elric 5y ago> you're almost always going to have some sort of singleton object somewhere in your application which just has these three pools, pre-configured for use I'm bemused by this statement, and I can't figure out whether this is an assertion rooted in supreme confidence, or just idle, wishful thinking. That being said, giving threading advice in a virtualized and containerized world is tricky. And while these three categories seem sensible, mapping the functions of any non-trivial system onto them is going to be difficult, unless the system was specifically designed around it.
- WatchDog 5y agoIf your app is fully non-blocking, doesn't it make sense to just do everything on the one pool, CPU bound tasks and IO polling. Rather than passing messages between threads.
- tadfisher 5y ago"Fully non-blocking" means "does no work". Ignoring the process' spawning thread, if your app performs CPU-bound tasks on a bounded thread pool, you will be leaving I/O throughput on the table as the number of tasks increases, since I/O-bound tasks will block on waiting for a thread.