5 ms·
> Can someone give a good argument of why subinterpreters are an interesting or useful solution to concurrency in Python? I will give it a shot. Subinterprete
by celeritascelery 3y ago
> Can someone give a good argument of why subinterpreters are an interesting or useful solution to concurrency in Python?
I will give it a shot.
Subinterpreters are better than multiple processes because:
- they have significantly less memory overhead
- they can move objects much faster between subinterpreters because they don’t need to serialize through a format like JSON
- since they are all in the same process you can implement things like atomics or channels easily.
Subinterpreter are better then no-Gil because:
- they make the code easier to reason about and debug relative to raw multi-theading
- they don’t negatively impact single threaded (basically all existing) python code performance
- they don’t require any changes to the C interface, preventing a fractured ecosystem
- they can’t have data races
- Spivak 3y ago> - they don’t negatively impact single threaded (basically all existing) python code performance I think this deserves an extra callout because even your multi-threaded Python programs are effectively single threaded and benefiting from the performance gain.
- deleted 3y ago[deleted]
- spacechild1 3y ago> - they can move objects much faster between subinterpreters because they don’t need to serialize through a format like JSON Why do you think that you would need to serialize to JSON? Pipes and sockets can deal with binary data just fine. With shared memory, there wouldn't be any difference at all. > - since they are all in the same process you can implement things like atomics or channels easily. This is also possible with shared memory. AFAICT the advantage of subinterpreters over subprocesses are: - lower memory overhead - faster creation/destruction time - ability to share global data (with subprocesses the data would either need to be duplicated or live in shared memory)
- celeritascelery 3y agoSure, you could do something like that. But a shared memory segment python with a stable binary object format doesn't exist (and isn't even being worked on). Comparing the proposed PEP 554 solution to a non-existent theoretical solution isn't very useful. But you do bring up some good points for ways you could achieve similar goals without the need to make the interpreters thread safe.
- spacechild1 3y ago> But a shared memory segment python with a stable binary object format doesn't exist There is multiprocessing.Queue (https://docs.python.org/3/library/multiprocessing.html#multiprocessing.Queue https://docs.python.org/3/library/multiprocessing.html#multi...). I don't know if it uses shared memory, or rather sockets or pipes, but this is just an implementation detail. My point is that there is no fundamental difference between isolated interpreters and processes when it comes to data sharing. Either way, you need a (binary) serialization format and some thread/process-safe queue. I would have naively assumed that you could repurpose multiprocessing.Queue for passing data between multiple interpreters; you would just need to replace the underlying communication mechanism (sockets, pipes, shared memory, whatever it is) with a queue + mutex. But then again, I'm not familiar with the actual code base. If there are any complications that I didn't take into acccount, I would be curious to hear about them. Interestingly, the PEP authors currently don't propose an API for exchanging data and instead suggest using raw pipes: > https://peps.python.org/pep-0554/#api-for-sharing-data https://peps.python.org/pep-0554/#api-for-sharing-data Of course, this is just a temporary hack. It would be ridiculous to use actual pipes for sharing data within the same process...
- sergiomattei 3y ago> they can’t have data races How so? Asking out of curiosity.
- celeritascelery 3y agoBecause they don't share memory. The C interpreters share memory, so they could have data races, but the python code can't. Just like how the C interpreter can have memory unsafely but python can't (or shouldn't).
- gyrovagueGeist 3y agoYour 'application' that uses them can absolutely have data races (you also have to fight the GC on every interpreter not just one). I think what they mean is that a local object within an interpreter without a shared (replicated) state does not have data races. Once you start message passing of course this is no longer true. Which isn't really different than a similar pattern of isolation in threading or multiprocessing.
- ptx 3y ago> they can move objects much faster between subinterpreters because they don’t need to serialize through a format like JSON That would be a huge advantage, but it's not there yet. According to PEP 554 [1] the only mechanism for sharing data is sending bytes through OS pipes, which is exactly the same as for multiprocessing and requires the same sort of serialization. [1] https://peps.python.org/pep-0554/#api-for-sharing-data https://peps.python.org/pep-0554/#api-for-sharing-data
- semiquaver 3y agoIs the overhead of pickle eg as used in multiprocessing.Pipe() actually a limiting factor in most circumstances? https://docs.python.org/3/library/multiprocessing.html#multiprocessing.Pipe https://docs.python.org/3/library/multiprocessing.html#multi... In a message passing system there’s always going to need to be some form of serialization. I’ll wager that pickle is fast and flexible enough for most cases and for those that aren’t, using something like flatbuffers or capn proto in shared memory wouldn’t be too much of a lift to integrate. Although all of that has long been possible in a multiple-process architecture, so I’m also curious to know if there are any real advantages to subinterpreters. From this message [1] linked to from the PEP it sounds like the author once thought that object sharing was a possibility, but if it’s not there seem to be no real benefits over multiprocessing and one big downside (the GIL). Contrast with ruby’s Ractor system [2], which is similar to the subinterpreter concept but allows true parallelism within a single process by giving each ractor its own interpreter lock, along with a system for marking an object as immutable so it can be shared among ractors. [1] https://mail.python.org/pipermail/python-ideas/2017-September/047122.html https://mail.python.org/pipermail/python-ideas/2017-Septembe... [2] https://github.com/ruby/ruby/blob/master/doc/ractor.md https://github.com/ruby/ruby/blob/master/doc/ractor.md
- gyrovagueGeist 3y agoIn my experience, yes, pickle is painfully slow and can't be used for anything real. It's fantastic for prototyping (especially https://mpi4py.readthedocs.io/en/stable/reference/mpi4py.MPI.Pickle.html https://mpi4py.readthedocs.io/en/stable/reference/mpi4py.MPI...) but it will be your bottleneck. But I work in more computational science so I understand my constraints are different than most folks.
- kmod 3y agoCouple corrections: - They absolutely do have to serialize, usually via pickle. I'm pretty sure objects are not sharable between subinterpreters and there is not a plan for that. The main reason people think subinterpreters are good ("you can just share the memory!") is not actually true. - They don't require any changes to the C interface because those changes were already made, and a fair amount of cost was paid by C library maintainers. So it's true, subinterpreters are at an advantage in this regard, but that's more of a political question than a technical one
- kzrdude 3y agoEric Snow mentioned in his Pycon talk that memory sharing would be used, especially big data blobs, arrays etc. Sure, not directly sending python objects, but passing pointers can be done.
- kzrdude 3y agohttps://www.youtube.com/watch?v=3ywZjnjeAO4 https://www.youtube.com/watch?v=3ywZjnjeAO4 after 19 minutes