4 ms·
> On the JVM, native threads are really expensive, especially in terms of memory. I don't think this is true if you create have approximately 1 thread for each
by vemv 5y ago
> On the JVM, native threads are really expensive, especially in terms of memory.
I don't think this is true if you create have approximately 1 thread for each CPU core (which is the ideal scenario in a non-async app - why would you create more threads than CPUs?).
A JVM thread is basically an OS thread. So on a 64-core server you'd have 64 threads, give or take. While a server can handle a few thousands threads quite seamlessly.
- btbuilder 5y agoIf you are not using async IO and have a thread per core, if (as is quite common) your program needs to make IO calls (network being likely worst case) then the thread will sleep on a blocking IO call and CPU will be idle even if there’s work in a queue.
- vemv 5y agoIO doesn't cause the whole CPU to be idle. It simply causes a context switch, which will allow another thread to do work.
- kaba0 5y agoBut it might not come back to the thread that made the IO only “much” later when the IO call interrupts. So a given thread will spend more time than necessary inside the thread pool, stealing resources from other threads.
- btbuilder 5y agoIf you only have one thread per core it is possible there will be no runnable threads while there is still work in the queue. Other threads could be waiting for IO or already running.
- Skinney 5y ago> which is the ideal scenario in a non-async app - why would you create more threads than CPUs? For workloads that involves a lot of IO. If you perform much more than 2-3 IO ops per request, a lot of time spent per request is simply waiting for IO response, meaning your CPUs are mostly idle.
- vemv 5y agoThis seems the sort of problem that is well-solved by a pipeline pattern. Each member of the pipeline is a thread pool that does exactly one type of job. One can allocate thread pools sensibly relative to CPU count and achieve ideal task<->CPU allocation.
- lmm 5y agoIf your application is simple enough that the pipeline stages fit neatly into the number of CPU cores you have, that's great. But for nontrivial applications one threadpool per pipeline stage would generally mean many more threads than CPU cores, at which point you're doing a lot of wasteful context switching. And since pipelines are one-way (unlike a function stack where you can return to the caller) there's a significant maintainability cost to structuring your program like that.
- vemv 5y agoA large monolithic app with all sorts of workloads going on in it seems hard to reason about, measure, GC-tune, HA, etc in the first place. So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count.
- lmm 5y ago> So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. That makes the context switching problem worse. If you're running 100 services with 8 threads each on an 8 core CPU, that's got the same overhead problem as running 800 threads in a single service - potentially worse if the kernel has to do a more thorough context switch (e.g. TLB flush for spectre mitigation) when switching between separate processes. (And if you're going to solve that problem by running each service in its own container or VM then again that only makes the context switches heavier - fundamentally if you're running 800 threads on 8 physical cores then you've got to switch between them somehow). > It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count. Well no shit if you just buy more CPU cores whenever you want to run more code then you don't need to worry about getting decent performance out of your hardware. But most of us are in businesses that can't afford one physical CPU core per IO-performing function in your codebase, which is what you'd need to sustain that approach (not to mention perfectly predicting how much load every single function needs). I guess when you add the 129th function to your codebase you then throw away all your super-expensive servers and buy even more expensive 256-core replacements? Or you forcibly split each service at the 128-function boundary so that you can deploy it as two microservices on two machines?
- clhodapp 5y agoThe model you are describing (incoming requests wait on a queue until a thread is ready to work on them) is a very simple form of async processing. Essentially, you are saving your context in some (relatively) small objects on the heap and waiting for the right time for that context to get picked up and processed. It's only a slight logical extension to allow you to suspend and resume multiple times in handling a single request. As far as using a fixed pool goes: others have replied similarly but by doing that you are completely letting your CPU cores go to waste whenever you are waiting for any underlying IO. For many classes of applications, they will end up topping out at only a few percent CPU usage.
- lmm 5y ago> I don't think this is true if you create have approximately 1 thread for each CPU core (which is the ideal scenario in a non-async app - why would you create more threads than CPUs?). Right, so then how do you structure your program - particularly an I/O-bound program - to use that small pool of threads effectively? Imagine you're building a REST frontend that queries a couple of datastores and aggregates the results - you get a web request, fire off a request to some coordinator datastore, get a response back from that, fire off a bunch of requests to other datastores based on that, then form the results into some JSON and send them back to the client. If you do blocking requests then you waste most of your threads most of the time (they'll be idle waiting for responses from the datastores). If you use NIO then that solves that problem, but you've still got to get each thread to actually do the right kind of processing - you want each thread to run an event loop where it checks for returned results from the datastores, matches those up with the requests that they belong to, and does the next step of processing for that request whether it's firing off more requests to other the datastores or composing together the results and sending them back to the original client. And you've got to also handle timing out stale requests etc. Now you can write a server literally like that, with a global buffers for each possible request type that hold the state that you need to pick up handling that request again. But it's a nightmare to maintain, as you're effectively forced to write unstructured programs where you do a GOTO for each I/O operation - there's no easy way to trace through the handling of a single request or run part of your program for testing. So people generally prefer to use an async framework that will handle keeping track of each suspended request and what to do next when the result comes back, and multiplex those onto the native threads for you. That could be something where you call an I/O API and pass a callback to run when the result comes back (and the framework takes care of running the callback on some thread when the result is ready, without blocking a thread in the meantime), or something where there's a special operator to say "suspend this function here, package up the rest of it as a continuation, and continue with that continuation when the I/O result comes back". Or, as Loom is doing, it could be something similar that works completely invisibly whenever the programmer calls particular functions. There are tradeoffs to all these approaches to async, but any of them is going to perform a lot better than the naive approach of just using a native thread for each request, and they're all (hopefully) going to be more maintainable than manually writing an event loop that does the right thing each time.