14 ms·
Genuine question - is there a concrete and practical problem that can't be solved without Loom green threads on the JVM? Generally I think that programs can be
by vemv 5y ago
Genuine question - is there a concrete and practical problem that can't be solved without Loom green threads on the JVM?
Generally I think that programs can be structured nicely and have them perform optimally by using traditional threads, backed by idiomatic usage of the java.util.concurrent package and integration of those vanilla threads with the NIO package.
IOW, using NIO doesn't inevitably imply using an async framework (be it bespoke, or official like Loom)?
Personally I've always had my reservations around async, it seems a paradigm that makes things gratuitously harder.
- clhodapp 5y agoOn the JVM, native threads are really expensive, especially in terms of memory. Loom comes from the same place of skepticism of async that you are expressing. Its goal is to make the imperative, blocking style as efficient as the async style while retaining familiar debugging tooling. Personally, I'm more skeptical of imperative, blocking programming but Loom does seem like it can only be a good thing to have access to on the JVM platform.
- vemv 5y ago> On the JVM, native threads are really expensive, especially in terms of memory. I don't think this is true if you create have approximately 1 thread for each CPU core (which is the ideal scenario in a non-async app - why would you create more threads than CPUs?). A JVM thread is basically an OS thread. So on a 64-core server you'd have 64 threads, give or take. While a server can handle a few thousands threads quite seamlessly.
- btbuilder 5y agoIf you are not using async IO and have a thread per core, if (as is quite common) your program needs to make IO calls (network being likely worst case) then the thread will sleep on a blocking IO call and CPU will be idle even if there’s work in a queue.
- vemv 5y agoIO doesn't cause the whole CPU to be idle. It simply causes a context switch, which will allow another thread to do work.
- kaba0 5y agoBut it might not come back to the thread that made the IO only “much” later when the IO call interrupts. So a given thread will spend more time than necessary inside the thread pool, stealing resources from other threads.
- btbuilder 5y agoIf you only have one thread per core it is possible there will be no runnable threads while there is still work in the queue. Other threads could be waiting for IO or already running.
- Skinney 5y ago> which is the ideal scenario in a non-async app - why would you create more threads than CPUs? For workloads that involves a lot of IO. If you perform much more than 2-3 IO ops per request, a lot of time spent per request is simply waiting for IO response, meaning your CPUs are mostly idle.
- vemv 5y agoThis seems the sort of problem that is well-solved by a pipeline pattern. Each member of the pipeline is a thread pool that does exactly one type of job. One can allocate thread pools sensibly relative to CPU count and achieve ideal task<->CPU allocation.
- lmm 5y agoIf your application is simple enough that the pipeline stages fit neatly into the number of CPU cores you have, that's great. But for nontrivial applications one threadpool per pipeline stage would generally mean many more threads than CPU cores, at which point you're doing a lot of wasteful context switching. And since pipelines are one-way (unlike a function stack where you can return to the caller) there's a significant maintainability cost to structuring your program like that.
- vemv 5y agoA large monolithic app with all sorts of workloads going on in it seems hard to reason about, measure, GC-tune, HA, etc in the first place. So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count.
- lmm 5y ago> So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. That makes the context switching problem worse. If you're running 100 services with 8 threads each on an 8 core CPU, that's got the same overhead problem as running 800 threads in a single service - potentially worse if the kernel has to do a more thorough context switch (e.g. TLB flush for spectre mitigation) when switching between separate processes. (And if you're going to solve that problem by running each service in its own container or VM then again that only makes the context switches heavier - fundamentally if you're running 800 threads on 8 physical cores then you've got to switch between them somehow). > It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count. Well no shit if you just buy more CPU cores whenever you want to run more code then you don't need to worry about getting decent performance out of your hardware. But most of us are in businesses that can't afford one physical CPU core per IO-performing function in your codebase, which is what you'd need to sustain that approach (not to mention perfectly predicting how much load every single function needs). I guess when you add the 129th function to your codebase you then throw away all your super-expensive servers and buy even more expensive 256-core replacements? Or you forcibly split each service at the 128-function boundary so that you can deploy it as two microservices on two machines?
- clhodapp 5y agoThe model you are describing (incoming requests wait on a queue until a thread is ready to work on them) is a very simple form of async processing. Essentially, you are saving your context in some (relatively) small objects on the heap and waiting for the right time for that context to get picked up and processed. It's only a slight logical extension to allow you to suspend and resume multiple times in handling a single request. As far as using a fixed pool goes: others have replied similarly but by doing that you are completely letting your CPU cores go to waste whenever you are waiting for any underlying IO. For many classes of applications, they will end up topping out at only a few percent CPU usage.
- lmm 5y ago> I don't think this is true if you create have approximately 1 thread for each CPU core (which is the ideal scenario in a non-async app - why would you create more threads than CPUs?). Right, so then how do you structure your program - particularly an I/O-bound program - to use that small pool of threads effectively? Imagine you're building a REST frontend that queries a couple of datastores and aggregates the results - you get a web request, fire off a request to some coordinator datastore, get a response back from that, fire off a bunch of requests to other datastores based on that, then form the results into some JSON and send them back to the client. If you do blocking requests then you waste most of your threads most of the time (they'll be idle waiting for responses from the datastores). If you use NIO then that solves that problem, but you've still got to get each thread to actually do the right kind of processing - you want each thread to run an event loop where it checks for returned results from the datastores, matches those up with the requests that they belong to, and does the next step of processing for that request whether it's firing off more requests to other the datastores or composing together the results and sending them back to the original client. And you've got to also handle timing out stale requests etc. Now you can write a server literally like that, with a global buffers for each possible request type that hold the state that you need to pick up handling that request again. But it's a nightmare to maintain, as you're effectively forced to write unstructured programs where you do a GOTO for each I/O operation - there's no easy way to trace through the handling of a single request or run part of your program for testing. So people generally prefer to use an async framework that will handle keeping track of each suspended request and what to do next when the result comes back, and multiplex those onto the native threads for you. That could be something where you call an I/O API and pass a callback to run when the result comes back (and the framework takes care of running the callback on some thread when the result is ready, without blocking a thread in the meantime), or something where there's a special operator to say "suspend this function here, package up the rest of it as a continuation, and continue with that continuation when the I/O result comes back". Or, as Loom is doing, it could be something similar that works completely invisibly whenever the programmer calls particular functions. There are tradeoffs to all these approaches to async, but any of them is going to perform a lot better than the naive approach of just using a native thread for each request, and they're all (hopefully) going to be more maintainable than manually writing an event loop that does the right thing each time.
- AlisdairO 5y ago> On the JVM, native threads are really expensive, especially in terms of memory. If you're running on Linux, this is rather less the case. People see 1MB Xss values and assume that 1MB of 'real' memory is required to spawn a thread. In reality, on a standard linux system with overcommit enabled, stack pages are only consumed as they're written to - the rest is just virtual. Possibly you're talking about something else, but I commonly see this raised as a concern about creating threads in Java.
- pron 5y ago> In reality, on a standard linux system with overcommit enabled, stack pages are only consumed as they're written to - the rest is just virtual. 1. Once stack memory is committed, it is never uncommitted, and creating a new thread is relatively expensive. 2. Committing virtual memory is done at a page granularity. These make threads far too costly to serve as a construct for any single concurrent task.
- AlisdairO 5y ago...any very short lived concurrent task. A thread costs, what, 10 or 20us to create on a typical linux/x86 system? Perhaps it's more on the JVM? Look, I'm looking forward to Loom as much as anyone. It is clearly going to be a big leap forward to the java ecosystem, and I'm looking forward to some threaded model being appropriate for any job rather than just most of them. It's just I see this terror of creating a thread or two out in the wider world and it seems a touch overwrought to me.
- pron 5y ago10 or 20us can be longer than some IO operations. But the direct issue with OS threads is simply that you can't realistically have too many of them, and so they limit your level of concurrency. If the number of threads you can maintain is more than the average request rate times the average latency of each transaction (if it were processed on a single thread), then you're all good; if it's less, then your server will crash as requests pile up. Fear doesn't play a role here, just Little's law.
- papercrane 5y ago> Genuine question - is there a concrete and practical problem that can't be solved without Loom green threads on the JVM? I think that's too high a bar. I think the right question should be "will Loom make it easier to write well performing and maintainable solutions to real problems?" I think by making threads a much cheaper resource the answer will be yes. For example, the simplest way to write a network server that can handle more than one client at a time is to fork a new thread for every client. Using that model you quickly run into a wall with OS threads in Java, and if you want more performance need to look into using NIO and event loops. With virtual threads you'll be able to go much further with the thread-per-client model.
- vemv 5y ago> For example, the simplest way to write a network server that can handle more than one client at a time is to fork a new thread for every client. This seems much debatable. For me the simplest way is creating a thread-per-core pool, and have clients integrate with said pool via a queue. It's a well-understood paradigm that has worked for a long time. I get the sense of abstraction of having a "thread" per client, but is it really the simplest choice if it needed creating a multi-year project like Loom first? A precedent that comes to mind is Clojure core.async. It pursues the same sense of abstraction. But does one gain much if the implementation has its own quirks, flaws, etc? Any given abstraction is one bug away from leaking. So, it's well plausible that the same can happen with Loom - more stuff to learn, more underlying complexity, bigger surface area for bugs. It all seems too gratuitous too me, so my high-bar question/questioning remains :)
- gravypod 5y agoAll implementations of in-process thread multiplexing is sort of like a garbage collector. You have a resource, you manage the allocation of a resource to something that consume it, etc. A lot of house keeping and some overhead code that needs to execute before any work is done (each time). In my mind, it makes sense to have a single, well tested, and abstracted implementation that all other systems can use. You can see a need for this at least in the Android space: RxJava, Kotlin coroutines, etc. Also, in a user-managed thread multiplexing system you can't guarantee "fair" scheduling. If you have some DOS bugs in your server you can lock up an entire serving thread. A language-level multiplexing system will not have these same kinds of edge cases. There are many places this sort of thing becomes helpful. We could also see concurrent processing utilities that could be very helpful. See this for more info on how this could lead to cleaner programs: https://www.youtube.com/watch?v=f6kdp27TYZs https://www.youtube.com/watch?v=f6kdp27TYZs
- tybit 5y ago> Personally I've always had my reservations around async, it seems a paradigm that makes things gratuitously harder. This is one of the major driving forces behind Loom being done so differently to how they’re done elsewhere. Async became so popular due to sync code not scaling for large concurrent use cases. Loom is about leaving sync application code be, and optimising the underlying runtime instead.
- jayd16 5y agoI don't think there's any day to day problem that needs more threads than you can get ahold of. This is especially true when you consider you can tune the JVM accordingly for the special cases. That said, Loom is a neat tool and I expect new, more convenient patterns will pop up from Loom. I'll be interested to see how Java tackles OS thread problems after Loom. We'll still live in a world where most(all?) GUIs frameworks revolve around single thread concurrency and GPU access is still single threaded. I don't see how the implicit threading of loom is going to work for that, and it also seems like a line in the sand has been drawn in the face of async/await. We'll just have to see what comes of it.
- kaba0 5y agoThe JVM will still create an OS thread by the default way. You have to explicitly create virtual threads to use loom.
- pjmlp 5y agoUnless they changed the behaviour, that wasn't the case last time I checked a talk on it. Rather virtual threads get created by default and an explicit call, or thread factory can be used to get the value back. At least that is how I remember it.
- jayd16 5y agoYeah, so if you have to explicitly call the new API when you need to use it form an OS thread, its not very implicit. That's what makes me curious about how they'll smooth that out.
- kaba0 5y agoIf you already have an executorservice, calling createVirtualThread is all you need to use Loom afaik.
- sriram_malhar 5y agoWe can agree that linearly laid out blocking code is simpler to write and debug than async code with callbacks. Ideally, one shouldn't even have to annotate functions as async. If an API call blocks, let it block. That would simplify a lot of conditional workflows immensely. The trouble is that threads are heavyweight, particularly in context switching. So the user has to work around it by explicitly creating async functions, task schedulers, futures and so on, to avoid having to block a crucial thread (as in GUIs or network event loops). Languages like go and erlang help you avoid that by having the entire ecosystem (libraries included) give you permission to write blocking code, and that under the hood, they would take care of fast context switching and async I/O. That is what loom hopes to do. Ideally in this world, the ordinary developer would not have to use nio at all ... simple blocking APIs should suffice.