4 ms·
Process bottlenecks are a design problem, not a language or syntax problem; and are mitigated largely by a few points that can be factored in during design or P
by bitwalker 6y ago
Process bottlenecks are a design problem, not a language or syntax problem; and are mitigated largely by a few points that can be factored in during design or PR review:
- Be wary of places where you have N:1 process dependencies, where N is large and the number of messages exchanged between each member of N and the single process are frequent/numerous. Since each process can only handle received messages sequentially, there is little point in spawning a lot of tasks in parallel if each process has to talk to the same upstream process to do anything
- Set up telemetry that samples the number of messages sitting in the process mailbox; if a process is becoming a bottleneck, it is going to be frequently overloaded and have a lot of messages in its mailbox. If you have the telemetry, you can see when this starts to happen, and take steps to deal with it before it starts causing problems for you. Likewise, its probably useful in general to have telemetry on how long each unit of work takes in server-like processes, so you can get a sense of throughput and factor that data into your design.
- Avoid sending large messages between processes, instead spawn a process to hold the data and then send a function to that process which operates on the data and returns only the result; or store the data in ETS if you have a lot of concurrent consumers. It can also be helpful to denormalize the data when you store it in ETS so you can access specific parts of it without copying the entire object out of ETS on every access. The goal here is to make messaging cheap and avoid copying lots of data around.
- Take steps to ensure process dependencies in your design are structured as trees, i.e. avoid dependency graphs that can contain cycles. It is all too easy for a change to introduce the possibility of deadlock if you play fast and loose with what processes can talk to each other. If your process dependencies mirror your supervisor tree, then you can protect against this by only allowing dependencies between branches in one direction (usually toward the parts of the tree that were started earlier in the supervisor tree)
I think the problem is that Elixir is still relatively young, and due to the language evolution and the lack of established documented doctrine from the Erlang community, there is a lot of techniques, tips, design patterns, etc., that are being rediscovered; likewise there are a lot of seemingly good ideas that turn out to be not so great in practice, but are encountered on the road to the truly sound patterns. So you get a lot of people writing about the lessons they are learning, and because of the gaps in knowledge, the result is that the information may be missing things, or providing a more complex solution when there is a simpler one, etc. Ultimately this is an important process, and now that Elixir has largely stabilized, this will only improve (and its is already pretty good, certainly far better than when I first started with the language years ago).
- nickjj 6y agoThanks. Do you have any code examples or practical applications on how to apply most of those bullet points? Those are all very daunting things to approach without examples but they sound very important. In Python or Ruby I would have just throw things into Redis as needed without thinking about it and this hasn't failed yet in years of development time with tens of millions of events processed (over time). Send ID of DB object to worker, let the worker library deal with it, look up the ID in the DB when the work is currently being done, let worker library deal with the rest and move onto the next thing. And for caching, it's just a matter of decorating a function or wrapping some lines of code to say it should be cached and everything works the same with 1 or 10 web servers when the state is saved into Redis (major web frameworks in Python and Ruby support this with 1 line of configuration).
- bitwalker 6y agoFor the use case you are describing, none of my points are important really - an HTTP request that hits a database, then pushes something onto a queue for background processing doesn't exhibit any problems from a process bottleneck point of view on that end of things. You still need to have some logic to deal with backpressure from the queue, but that is a language agnostic concern. Where you could hit a bottleneck might be in the background processing though, take for example the following scenario: - A pool of N background job worker processes each pull an item off a queue, and spawn a process to perform the task in isolation - A singleton process S provides exclusive access to some resource - Each task calls some code which needs to interact with the resource controlled by S. The problem with the above is that all of that concurrency/parallelism is nullified by the fact that the tasks are all going to block on S to do their work, the bottleneck of the design. To be clear, you should always gather telemetry first, but lets assume that you've gathered that and you can clearly see that this bottleneck is an issue (the process mailbox has frequently got many messages waiting to be received, the average time to completion for jobs is increasing). To solve this depends on why the resource is held by S in the first place. If its because the resource is not thread safe and requires exclusive access, then unless you can find a way to avoid needing the resource in every task, there isn't much you can do, but this should be fairly uncommon in practice. If S exists because you needed to store some shared state somewhere, and someone told you that an Agent or GenServer was the way to go, then you could move that data to ETS and make it publically accessible so that functions which operate on that data read it from ETS directly rather than call the process. Now you've removed that bottleneck entirely. If S exists because it needs to protect access to some data, but not all of it, and most tasks don't need to access the protected data, then you can move the parts that do not need to be protected into ETS, and keep the rest in the process. This might reduce the amount of contention on that singleton process by a huge amount, but if even half the processes no longer need to block on accessing it, then you've regained at least that much concurrency in the task processing code. --- The example above is something I've seen numerous times, but the important pattern to note is that you have some task that you've tried to parallelize by spawning multiple processes, but that task itself depends on something that is not, or cannot be done concurrently/in parallel. Any time this pattern arises, you need to either find a way to enable concurrency in that dependency, or you should avoid doing the task in parallel in the first place. This is ultimately true of any parallelizable task - its only parallelizable if all of the tasks dependencies are themselves parallelizable, otherwise you end up bottlenecked on those dependencies and you've gained little to no benefit. Where it becomes a bigger problem is when you consider the system at a higher level. Bottlenecks reduce throughput, which may end up, via backpressure, causing errors on the client due to overload, or depending on the domain, data being dropped because it can't be handled in time (e.g. soft real-time systems). I don't have any code examples that really encompass all of this in one place, if you are interested in something specific, I can try to throw something together for you. Or if you have specific questions I can point you to some resources I've used to help understand some of these concepts.