4 ms·
In general, what are the advantage of a thread-per-request model? Better load balancing between cores?
by xoranth 3y ago
In general, what are the advantage of a thread-per-request model? Better load balancing between cores?
- onetimeuse92304 3y agoThread per request can never have better load balancing between cores than a well designed, custom solution. You are essentially asking the operating system to do the scheduling for you. But the OS will never be able to do it perfectly as it has no knowledge of what your application is doing. The main advantage of OS scheduling is that you get pretty good results without having to think about it at all. Pretty good, but never perfect.
- xoranth 3y ago> You are essentially asking the operating system to do the scheduling for you. But the OS will never be able to do it perfectly as it has no knowledge of what your application is doing. From LMAX presentations, it looks like they want you to split your application into tasks [1], define a graph of task dependencies, have each core process a particular kind of task and have task processors communicate their producers via a ring buffer. In particular, the allocation of tasks is static. The use of a ring buffer means that there is very little contention, and task processing is very efficient, but some cores might end up underutilized. On the other hand, if you have a thread per request, and allow them to migrate between cores, idle cores can steal tasks from busy ones. So in theory you could get better utilization, but task processing is less efficient since you need to share more data between cores. ("threads" don't need to be OS threads, they can be green threads) That said, I am not sure if GP meant this by thread-per-request, or "legacy" applications that use a thread pool, or something else. [1]: https://www.slideshare.net/trishagee/introduction-to-the-disruptor#39 https://www.slideshare.net/trishagee/introduction-to-the-dis...
- onetimeuse92304 3y agoI know all about LMAX architecture, at least all that has been published (see my other comments for this submission). Static allocation is a special case of scheduling. You decide which parts of the process run on which core -- the scheduling in this case is done at design or configuration time. > On the other hand, if you have a thread per request, and allow them to migrate between cores, idle cores can steal tasks from busy ones. So in theory you could get better utilization, Migrating your tasks between cores is nothing you can't design into your application. For example, in a typical event driven architecture where you have worker threads running each on separate core, there would be something to decide where the task is queued and usually the logic will take into account how busy particular worker thread is. The operating system does nothing in this case, what it sees is a number of threads each running on its separate core that don't need to be preempted (hopefully).
- xoranth 3y ago> For example, in a typical event driven architecture where you have worker threads running each on separate core, there would be something to decide where the task is queued and usually the logic will take into account how busy particular worker thread is. Wouldn't that work well only if the time taken by each task is predictable? I.e. you mention working on a trading system. But in a trading system you want to run the same branchless code path regardless of the kind of incoming event and whether you are sending an order or not after running the trading logic. So the individual "task" is very predictable. On the other hand think of a task like "return all the comments for a certain page". The time taken by an individual task is unpredictable, proportional to the number of comments. So you'll regularly get one core getting enqueued a bunch of tasks with no comments, finishing quickly and then staying idle. With work stealing, after finishing, that core would get a chance at "stealing" tasks from other threads' queues. (of course, the architecture I am describing would be awful for a trading system) > The operating system does nothing in this case, what it sees is a number of threads each running on its separate core that don't need to be preempted (hopefully). Btw, I agree that pinning OS threads to each core and then layering something of your own on top of it is going to be faster. It is just that you can layer on top a green thread system (like Go), and get something thread-for-request -like.
- lowbloodsugar 3y agoI feel like you are stuck on the idea of “this should be simple for me” which, of course, favors the OS threads solution. If your point is that LMAX inst just a drop in substitute for OS threads then we agree. If your point is that OS threads produce better results than a thought out LMAX solution then we do not. Most people and organizations probably don’t have the need or the skill for LMAX anyway.
- citrin_ru 3y agoEasy to implement.
- nine_k 3y agoBut does anyone run an OS thread per request unironically? I thought that nearly every request-response server implementation would use a thread pool. The best, like Erlang, can give you the feeling of arbitrarily many extremely cheap threads, while also running on a thread pool.
- marginalia_nu 3y agoAs a devil's advocate argument, if you're doing serverside rendering, and basically getting 1 request per visit, sure there's overhead, but even like a landslide HN death hug is only a handful requests per second. A raspberry pi could feasibly serve that traffic spawning 1 thread per request. ... not that I think anyone is doing this outside of maybe some hobbyist building their own HTTP server for fun.
- gmfawcett 3y ago> does anyone run an OS thread per request unironically? Of course they do. There are loads of appropriate applications. Heck, people still run CGI programs unironically.
- sgift 3y agoAnd if your use case allows it it is a great model. Easy to setup, easy to debug, easy to run. Also, why I love the new virtual threads in Java. Will they work as good as hoped? No idea. Probably not for all use cases. But the direction to say: You know what, threads are great to program and debug compared to the alternatives, maybe let's find a way to make their performance better instead of putting up with async , is so refreshing.
- yencabulator 3y agoAs far as comparisons to thread-per-core go, thread per request applies whether it's an OS thread or a green thread or a Rust async function compiled into a state machine. Anything that multiplexes per-request contexts into a lesser amount of cores(/OS threads) has the same trade-offs, the difference is more on the easy-vs-optimized spectrum. Thread-per-core with fixed workloads behaves differently than all of those. Here's an example difference: in thread-per-request, any global state can be accessed from "anywhere", and thus you end up with locks, reference counts, GC, and what not. In thread-per-core, global state is sharded across cores and never accessed "from the outside", and thus needs no locks/atomics (beyond the messaging primitive).