3 ms·
This seems the sort of problem that is well-solved by a pipeline pattern. Each member of the pipeline is a thread pool that does exactly one type of job. One c
by vemv 5y ago
This seems the sort of problem that is well-solved by a pipeline pattern. Each member of the pipeline is a thread pool that does exactly one type of job.
One can allocate thread pools sensibly relative to CPU count and achieve ideal task<->CPU allocation.
- lmm 5y agoIf your application is simple enough that the pipeline stages fit neatly into the number of CPU cores you have, that's great. But for nontrivial applications one threadpool per pipeline stage would generally mean many more threads than CPU cores, at which point you're doing a lot of wasteful context switching. And since pipelines are one-way (unlike a function stack where you can return to the caller) there's a significant maintainability cost to structuring your program like that.
- vemv 5y agoA large monolithic app with all sorts of workloads going on in it seems hard to reason about, measure, GC-tune, HA, etc in the first place. So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count.
- lmm 5y ago> So a modern choice would be either to have specialized background job processes, or microservices. Each such process would have manageable thread pools and a good CPU fit. That makes the context switching problem worse. If you're running 100 services with 8 threads each on an 8 core CPU, that's got the same overhead problem as running 800 threads in a single service - potentially worse if the kernel has to do a more thorough context switch (e.g. TLB flush for spectre mitigation) when switching between separate processes. (And if you're going to solve that problem by running each service in its own container or VM then again that only makes the context switches heavier - fundamentally if you're running 800 threads on 8 physical cores then you've got to switch between them somehow). > It also is a non-trivial fact that servers can perfectly have 128 cores or such, just like TB-sized memory isn't unheard of for JVMs. So that's another option for "less modern" apps: go the thread pool route, increase CPU count. Well no shit if you just buy more CPU cores whenever you want to run more code then you don't need to worry about getting decent performance out of your hardware. But most of us are in businesses that can't afford one physical CPU core per IO-performing function in your codebase, which is what you'd need to sustain that approach (not to mention perfectly predicting how much load every single function needs). I guess when you add the 129th function to your codebase you then throw away all your super-expensive servers and buy even more expensive 256-core replacements? Or you forcibly split each service at the 128-function boundary so that you can deploy it as two microservices on two machines?
- vemv 5y ago> And if you're going to solve that problem by running each service in its own container or VM then again that only makes the context switches heavier [...] Quite obviously I meant running a microservice / job processor in a different "node" (aws instance, whatever). It barely makes sense to let microservices compete with each other for CPU - one precisely seeks greater isolation. > Well no shit if you just buy more CPU cores whenever you want to run more code then you don't need to worry about getting decent performance out of your hardware. Large codebases relate strongly to large businesses that can afford hardware matching their scale. > I guess when you add the 129th function to your codebase you then throw away all your super-expensive servers and buy even more expensive 256-core replacements? Ok I regret spending time over your replies.
- lmm 5y ago> Quite obviously I meant running a microservice / job processor in a different "node" (aws instance, whatever). That's still just sweeping the problem under the carpet. Either you're getting a physical core for each function, in which case you're paying for lots of cores. Or you're multiplexing functions onto cores at some level (whether that's your own code, shared hosting, AWS lambda, or whatever) in which case someone (you or your service provider) is doing a bunch of context switches that have to be paid for somehow. > It barely makes sense to let microservices compete with each other for CPU - one precisely seeks greater isolation. For services big enough to need a full CPU (or a full core) all the time, you may as well give them a full CPU. But that implies a much coarser level of granularity than you've been suggesting. If your services really are "micro" then it makes a huge amount of sense to share a CPU between multiple services because most of them don't need a whole CPU the whole time. > Large codebases relate strongly to large businesses that can afford hardware matching their scale. The code grows faster than the hardware use. I've worked on codebases that were deployed on thousands of machines - but they had millions of lines of code and tens if not hundreds of thousands of functions. One of them even did take the approach of running multiple steps of the same logical computation on separate machines - but this was still done by having a yield point in the code and a scheduling framework that dispatched tasks onto grid compute machines, both because even though the code gets compiled into a state machine representation (i.e. one-way flows, de facto unstructured programming - the same as your pipeline approach) you don't want to have to maintain it in that form, and even if you did have code that looks like that, manually deciding how many cores to allocate to processing each state would not have been tractable. > Ok I regret spending time over your replies. Look, I honestly thought you were doing some kind of elaborate troll. What on earth kind of environment are you working in where you have more cores than code?
- kaba0 5y agoThe CPU’s scheduler has much less information available than the JVM does. A blocking call can be interpreted as non-blocking in case of a runtime, meaning more efficient CPU-usage.
- Skinney 5y agoWhat you're proposing is more work than simply spawning one thread per task, which is what Loom let's you do at scale. Why change your app into microservices if you don't need to?