7 ms·
Do you happen to have any papers about the current efforts to remove the GIL? I love Python, and use it a lot for ETL type work, but if threading worked well,
by hermitdev 10y ago
Do you happen to have any papers about the current efforts to remove the GIL?
I love Python, and use it a lot for ETL type work, but if threading worked well, I could/would possibly use it for far more purposes.
- brianwawok 10y agoCan you give an example where the GIL is really holding you back? Because with multiprocessing and greenlets, 99.99% of concurrency problems are trivilially solved by current Cython.
- deleted 10y ago[deleted]
- mrits 10y ago2 TB hash join
- doubleunplussed 10y agoIsn't that up to the RDBMS whether than's multithreaded or not? Unless the RDBMS is implemented in Python, CPython doesn't force extension code to be single-threaded. Just Python bytecode.
- foo101 10y agoThat's pretty much the point, isn't it? If I need true multithreading, then I am forced to write an extension in C or offload the multithreading work to another process (such as RDBMS in your case). It would have been nice if true multithreading was possible in Python itself. It would immediately make Python more useful in a variety of scenarios where splitting the work into multiple processes is not optimal or more convoluted.
- m_mueller 10y agoExactly, see my own post a few branches up.
- doubleunplussed 10y agoSure, I guess I'm just used to writing my performance bottlenecks in a lower level language already, so I'm used to the GIL not actually being held most of the time in any intensive computation. So if I want to call two Fourier transform functions at the same time in Python I can, because neither of them is implemented in Python and so they don't hold the GIL. That's the kind of parallelism use case I most often see come up, so although the GIL dismayed me early on I've come to see it as pretty irrelevant. But maybe it makes more sense for other applications, for the performance critical parts of the code to be actual pure Python. I do mostly numerical simulations, so pure Python is usually a non-starter, you fix that long before you think about parallelism.
- mrits 10y agoIf your answer to any GIL's is to write in a different language then I guess you don't have a problem.
- doubleunplussed 10y agoIt's not that I write the whole program in another language, it's that I either write the bottleneck in another language (usually Cython), or it turns out that the Python package I'm calling already has its bottlenecks written in another language, whether I wrote it or not. Day to day, I'm writing Python code which is actually parallel because a large fraction of the run time is dominated by the by things that aren't pure Python. I suspect this is true even for people who are not going out of their way to make it true. It's simply the case that most RDBM systems, Fourier transforms, etc with Python bindings are not written in Python. The GIL sounds scary, but I think people overestimate the fraction of time it is actually held in their code.
- m_mueller 10y agoactually GP, but it has held me back in the past. I'm writing a transpiler that uses global information from codebases, and so it transpiles potentially hundreds of files at once and creates rather complex data structures. Compute bound for quite a while, so I tried speeding it up with multiprocessing (since multithreading would be useless). But with multiprocessing it took longer to serialize/deserialize the complex datastructures for each process, so I had to give up. Next time I have time for this I'd probably try to use Jython as a drop-in replacement and see whether I can get it to run with GIL-less multithreading.
- orf 10y agoIt sounds like you have a couple of hot paths and are not optimizing them. I can't tell for sure without seeing any code but nothing in your post screams out "this will be slow" or "I need parallism/concurrency". Perhaps it's the data structures you are using?
- m_mueller 10y agoI already did extensive profiling and performance improvements, at this point I'm quite sure that if I could do multithreading on my lab's 24 core Xeon Haswell machines I'd be getting a nice speedup.
- xapata 10y agoSounds like you might be iterating dictionaries. That's much faster in Python 3.6 due to the compaction of dict storage.
- pfranz 10y agoLarry Hastings - Removing Python's GIL: The Gilectomy - PyCon 2016 https://youtu.be/P3AyI_u66Bw https://youtu.be/P3AyI_u66Bw (I'm pretty sure this is the video I'm thinking of) It's 30m, but worth it if you're interested. Not sure what progress has been made since then.
- esaym 10y agoPlease remember that threads are not the only way. If you can simply break your function/routine into a smaller piece that is independent, you can easy get by with a fork. (well, unless you are on windows..)
- hermitdev 10y agoBeen a since I tried to use the multiprocessing module. But, last time I did try, I ran into issues with it interacting poorly with pyodbc. It's been years, so I don't recall what the problem was, but I spent a few days trying to resolve or work around the issue with no satisfaction. Also, most of my Python scripts run on both Linux and Windows, so I have that restriction, as well.
- dagw 10y agohttp://pyparallel.org http://pyparallel.org is one of the more interesting experiments currently going on in the GIL area . They're basically working on removing all the practical limitations of the GIL without actually removing the entire GIL.
- trentnelson 10y agoPyParallel v1 was a nice checkpoint. I'm working on the next incarnation of it now.
- dagw 10y agoLooking forward to it. PyParallel is one of the more exciting python implementations out there