4 ms·
>Nothing unexpected there. Amdahl's Law in its glory. That's not Amdahl's Law. Amdahl's Law describes the expected total speedup from a speedup to a segment of
by adfgiknio 3y ago
>Nothing unexpected there. Amdahl's Law in its glory.
That's not Amdahl's Law. Amdahl's Law describes the expected total speedup from a speedup to a segment of the program. It is usually used in the context of adding parallelism but is more general than that. It does not assume parallelism has any overhead.
Amdahl's Law predicts performance will approach a limit with added cores, but still predicts that each core will improve performance.
>A fast running function finishes fast, and coordinating the job's execution over many cores requires doing work every time something finishes. If you split your job into chunks in the microsecond time range, then you'll be handling lots and lots of tiny minutia.
Not necessarily. Modern systems have pools with independent per-core queues and clever ways of distributing work to them with minimal overhead. Work-stealing pools don't have any overhead from synchronization except when a core runs out of work, so each core just runs a `for` loop most of the time. I have seen a simple (pseudocode) `list.parallelMap(foo)` give a linear speedup on a modern JVM even on a large host with a `foo` that takes <100ns.
Trying to manually batch work risks one thread running faster than the others and then idling when it runs out of work. This is common on modern hardware with heterogeneous cores and dynamic frequency scaling.
If the tasks require their own coördination or the runtime isn't smart, then manual batching may be needed. I wouldn't assume Python does the right thing.
- csdvrx 3y ago> Amdahl's Law predicts performance will approach a limit with added cores, but still predicts that each core will improve performance. And I would agree: on my systems, efficiency cores are used to handle other tasks and keep the power core cold and ready to use when needed (dynamic frequency scaling) > Trying to manually batch work risks one thread running faster than the others and then idling when it runs out of work. This is common on modern hardware with heterogeneous cores and dynamic frequency scaling. Yes, I think the issue here is more that python can't introspect the cgroup artificial limitations placed by docker on the CPU it's using, causing the kink past 9. If the goal is performance, I think it would be better to assign the cores manually + use cpu pinning + declare to the containers what resources it really has. On a NOHZ kernel, you can do manual assignments like nohz_full=1-3,5-7 rcu_nocbs=0-3,5-7 irqaffinity=4 This puts the IRQ burden on one core but you can do 2 (with =4,5) etc
- dale_glass 3y ago> That's not Amdahl's Law. Amdahl's Law describes the expected total speedup from a speedup to a segment of the program. It is usually used in the context of adding parallelism but is more general than that. It does not assume parallelism has any overhead. Yeah, it's a fiction. An useful one, but still a sort of spherical cow in a vacuum scenario. It doesn't exactly describe real hardware. > Amdahl's Law predicts performance will approach a limit with added cores, but still predicts that each core will improve performance. Because it's a fiction. You can have slowdowns with bigger tasks on real hardware, like if you run out of L1 cache. > Not necessarily. Modern systems have pools with independent per-core queues and clever ways of distributing work to them with minimal overhead. True, but that's Java and this is Python. I may be wrong, but since the GIL removal is very new still, I don't expect Python to have nearly the same level of efficiency in this regard. But I could be wrong somewhere, of course.
- Dylan16807 3y ago> Because it's a fiction. You can have slowdowns with bigger tasks on real hardware, like if you run out of L1 cache. Yes...? But the difference between it and reality is exactly why "optimal number of threads" is not Amdahl's Law.