6 ms·
Your criteria 2, 3, and 4 doesn't make much sense to me. We often have workloads that require multiple boxes, but we still want to make effective use of each bo
by colesbury 9y ago
Your criteria 2, 3, and 4 doesn't make much sense to me. We often have workloads that require multiple boxes, but we still want to make effective use of each box. Common server hardware has dozens of cores, which requires a lot of parallelism to fully utilize. The GIL hinders that, even when most of the work doesn't hold the GIL (see Amdahl's law)
Python multiprocessing doesn't work well with a lot of external libraries. For example, CUDA doesn't work across forks and many system resources can be shared across threads but not processes. Python objects must be pickled to be sent to another process, but not all objects can be pickled (including some built-in objects like tracebacks).
A lot of different parallel programming models can be built on top of threads (shared memory, fork-join, message passing), and to a certain extent they can be mixed. That's not true of Python multiprocessing, which only allows a narrow form of message passing. (It's also buggy, has internal race conditions, and easily leaks resources.)
The problem for CPython is that it may not be possible to remove the GIL without breaking the C API, and a lot of the benefit of Python is the huge number of high-quality packages, many of which use the C API.
- PrimHelios 9y agoCPython doesn't have any reservations about breaking the Python API between minor versions, so why care about the C API? I get where you're coming from, but they've already shown they don't care much for compatibility, so I don't see why that's a big obstacle.
- dom0 9y agoRemoving the GIL (in a non-braindead way) likely entails breaking all existing code using the C API. PyPy could do so without breaking cpyext, by maintaining the illusion of a GIL whenever control passes to cpyext.
- std_throwaway 9y agoDoes it lock the GIL so numpy can release it again immediately afterwards?
- dkersten 9y agoPerhaps it makes the unlock call a no-op before numpy tries to unlock it.
- PrimHelios 9y agoThat makes sense, I hadn't thought about the extent of the breakage.
- dom0 9y ago> (see Amdahl's law) Amdahl's law bears little relevance to throughput computing (i.e. most servers). > (It's also buggy, has internal race conditions, and easily leaks resources.) There is also at least one memory corruption bug in multiprocessing (linked a few months back by a fellow HN reader).