3 ms·
> You are essentially asking the operating system to do the scheduling for you. But the OS will never be able to do it perfectly as it has no knowledge of what
by xoranth 3y ago
> You are essentially asking the operating system to do the scheduling for you. But the OS will never be able to do it perfectly as it has no knowledge of what your application is doing.
From LMAX presentations, it looks like they want you to split your application into tasks [1], define a graph of task dependencies, have each core process a particular kind of task and have task processors communicate their producers via a ring buffer.
In particular, the allocation of tasks is static. The use of a ring buffer means that there is very little contention, and task processing is very efficient, but some cores might end up underutilized.
On the other hand, if you have a thread per request, and allow them to migrate between cores, idle cores can steal tasks from busy ones.
So in theory you could get better utilization, but task processing is less efficient since you need to share more data between cores.
("threads" don't need to be OS threads, they can be green threads)
That said, I am not sure if GP meant this by thread-per-request, or "legacy" applications that use a thread pool, or something else.
[1]: https://www.slideshare.net/trishagee/introduction-to-the-disruptor#39 https://www.slideshare.net/trishagee/introduction-to-the-dis...
- onetimeuse92304 3y agoI know all about LMAX architecture, at least all that has been published (see my other comments for this submission). Static allocation is a special case of scheduling. You decide which parts of the process run on which core -- the scheduling in this case is done at design or configuration time. > On the other hand, if you have a thread per request, and allow them to migrate between cores, idle cores can steal tasks from busy ones. So in theory you could get better utilization, Migrating your tasks between cores is nothing you can't design into your application. For example, in a typical event driven architecture where you have worker threads running each on separate core, there would be something to decide where the task is queued and usually the logic will take into account how busy particular worker thread is. The operating system does nothing in this case, what it sees is a number of threads each running on its separate core that don't need to be preempted (hopefully).
- xoranth 3y ago> For example, in a typical event driven architecture where you have worker threads running each on separate core, there would be something to decide where the task is queued and usually the logic will take into account how busy particular worker thread is. Wouldn't that work well only if the time taken by each task is predictable? I.e. you mention working on a trading system. But in a trading system you want to run the same branchless code path regardless of the kind of incoming event and whether you are sending an order or not after running the trading logic. So the individual "task" is very predictable. On the other hand think of a task like "return all the comments for a certain page". The time taken by an individual task is unpredictable, proportional to the number of comments. So you'll regularly get one core getting enqueued a bunch of tasks with no comments, finishing quickly and then staying idle. With work stealing, after finishing, that core would get a chance at "stealing" tasks from other threads' queues. (of course, the architecture I am describing would be awful for a trading system) > The operating system does nothing in this case, what it sees is a number of threads each running on its separate core that don't need to be preempted (hopefully). Btw, I agree that pinning OS threads to each core and then layering something of your own on top of it is going to be faster. It is just that you can layer on top a green thread system (like Go), and get something thread-for-request -like.
- lowbloodsugar 3y agoI feel like you are stuck on the idea of “this should be simple for me” which, of course, favors the OS threads solution. If your point is that LMAX inst just a drop in substitute for OS threads then we agree. If your point is that OS threads produce better results than a thought out LMAX solution then we do not. Most people and organizations probably don’t have the need or the skill for LMAX anyway.