4 ms·
Actor frameworks remove some of the ceremony around async message passing, but the central issue of scalability is "canonical state" and how well commands and q
by slver 5y ago
Actor frameworks remove some of the ceremony around async message passing, but the central issue of scalability is "canonical state" and how well commands and queries on that scale.
It's a bit like tech support in tiers. You have one CEO, but you don't get the CEO on every tech support call. You have bunch of layers trying to satisfy you until you get there.
The problem is those layers may also provide very poor UX.
- mhoad 5y agoI wish I was smart enough to get what you were saying here but I got a bit lost on the analogy. Could you ELI5 that for me? It sounded interesting
- rapsey 5y agoIt always comes down to your data that must be persistent and consistent. This data must be stored somewhere. The real problem is always storage. Message passing is all fine and good, but the real problem is your data. How scalable is your data layer in terms of size and operations per second.
- slver 5y agoWhen you scale, your basic solution is "I take this job done by one person/machine, and I split it to multiple people/machines". Sometimes you can't do that. Either because you're "outsourcing" part of the solution to someone beyond your control, or because the data relationships of your system require a "synchronizer", someone who works with this data in a serial, specific way. For example we have some government agency. Long queues, lots of waiting. One office for registration. What do we do? We add more registration offices, so people's forms can be processed in parallel. But every person needs to get a unique ticket number let's say. Now they still need to go through a single office to get this number, because the independent offices can't guarantee global uniqueness as they don't synchronize between each other. You need a "synchronizer" whose only purpose is to provide unique numbers. And this works, you have 100 offices and 1 "number" office. You scaled 100x over having one office for everything. But when you get to 200 offices, suddenly the queue in front of the "number" office starts to grow. You can't scale this office, because there's a "data relationship" in there that you can't split up and scale. You wanna identify and avoid such bottlenecks, and when you can't avoid them you want those bottlenecked systems to do as little as possible (so they can run as fast as possible) and to be hit as rarely as possible (by having someone in front take care of common cases: for example you can have a front-desk before a "number" office that checks you have all needed documents to get a number, before you get sent to get a number, now the "number" office has less to do). Sorry that's very abstract, but without specific examples it's hard to explain anything.
- devoutsalsa 5y agoThe concept was nice of the one office that gives out numbers is similar to auto incrementing integer IDs in DB. To get around one single thing having to generate all the numbers, you can have every office generate UUIDs, and toss in a timestamp if you need ordering. I’ve heard people are playing with UUID v6, too, but I’m not sold on it yet.
- mhoad 5y agoNo need for the apology, I appreciate you taking the time to event attempt explaining it :) Let me point out where I am falling over specifically because it might be more helpful. My understanding of Actors is that because they are essentially operating as in-memory objects with their own state, that with the exception of ad-hoc queries on my data (where SQL shines for example) I am doing the majority of my operations on various resources which handle persistence in a eventually consistent fashion and therefore aren't really obliged to immediately write to a database with each individual change in state thus freeing up the bottlenecks. This is all a pretty new concept to me so I feel like I might have made some bad assumptions in there for example, I am just not 100% sure where yet :)
- deleted 5y ago[deleted]
- mamcx 5y agoIn other words: You can't add many workers to a linear job and expect the task is done in SizeOfJob/Workers. Or: If you put many car lanes that go into one, you will get contention. This will manifest in many ways. But you "can" if along ALL the pipeline you have a way to coordinate the task!. However, if componentA is paradigmA and componentB is paradigmB you will get trouble. --- In practical term: Wanna kill your database? Open N connections launched by thousand of actors/async/threads. You need to instead use a pool, coordinate which jobs are fast to done or slow to complete. AND You need to architect the code to exploit both what this paradigm is, + what that make happy the DB (typical: NOT do sql + n, think in batches, group operations, etc) This is the part you don't get much help not matter what you use.
- patrec 5y ago> the central issue of scalability is "canonical state" Exactly. I've written a reasonable amount of Erlang, and yeah, from a purely technical perspective you'd be much better off handling concurrent requests with something actor based than async-du-jour dumpster fire (python is particularly bad). But the real difficulty ultimately is managing canonical persistent state, and actors won't magically solve that for you. Having said this, empirically speaking, almost none of the companies who believe they have scalability problems actually do and a relational DB combined with a not super inefficient architecture are all that's needed.
- samsquire 5y agoProblem with python is that async doesn't confer any performance benefit because of the GIL. You cannot actually run async functions in parallel. Each function makes forward progress independently but in serial. One problem I have is that the performance of computation is tied to performance of the storage, hence bringing computation to the storage for performance. If you have to communicate to get storage, then your solution will be slow.
- eloff 5y agoAsync is not really about parallel. Async in nodejs works the same way. The point is for IO to happen concurrently. To scale to multiple cores in Python or node you run multiple processes. If you're not sharing mutable state between them, that scales great. If you are, you really shouldn't be using those languages, they're not the right tool for the job.
- samsquire 5y agoI take advantage of I/O being parallel in Python in my mazzle continuous integration pipeline tool. I'm not sharing mutable state. I spin up a graph of python Threads and each joins others in a graph. This way we can run graphs in parallel. See this graph - the parts that look like this: dependency -> {parallel1; parallel2; parallel3} -> postparallel parallel1, parallel2, parallel3 can run in parallel in a separate python thread because the IO is parallel. postparallel joins parallel1, parallel2, parallel3 and waits for them all to complete. Where parallel1-3 is things like ansible, packer (slow), AMI builds, chef runs etc. https://github.com/samsquire/mazzle-starter/blob/master/architecture.png https://github.com/samsquire/mazzle-starter/blob/master/arch...