7 ms·
Exactly. I think a lot of the negativity about GIL comes from a misunderstanding about forking processes. If python is being used as a scripting language, and s
by commonlisp94 3y ago
Exactly. I think a lot of the negativity about GIL comes from a misunderstanding about forking processes. If python is being used as a scripting language, and spawning other tools, you're already getting free multi-core.
A similar misunderstanding exists about SQLite and concurrency.. but that's a topic for another time.
- AlphaSite 3y agoForking had a ton of its own downsides, it’s not a free lunch either, from poor ergonomics to communications overhead it works well for somethings and very poorly for others.
- miraculixx 3y agoYes, the same is true for free threading. Yet people assume free threading is free concurrency and that's the problem.
- usrbinbash 3y agoShow me someone who actually knows how threads work and what writing threading code entails who assumes that. I am perfectly aware that threads are not free. Just as I am perfectly aware that a context switch between threads is less expensive that switching a new process onto the core, and that IPC requires kernel involvement.
- commonlisp94 3y ago> context switch between threads is less expensive that switching a new process onto the core But the overhead of a context switch for a thread and process is very similar. The main difference is whether memory is shared by default.
- usrbinbash 3y ago> is very similar. Except that the threads share the exact same virtual address space, and processes do not, which makes the thread context switch faster. And that is to say nothing about the setup and teardown process, which for a process involves copy-on-demand'ing the entire memory, but for a thread merely setting up its own stack.
- commonlisp94 3y ago> Except that the threads share the exact same virtual address space, and processes do not, which makes the thread context switch faster That's what I said. But it's really not much. I'm afraid we will need numbers now to continue the conversation. If I measured would you be open to changing your opinion? Or are you committed to this topic, so that it would have no bearing?
- kerkeslager 3y agoWell, I don't know how you'd test that, but you should really consider testing the other half of the post you're responding to which you ignored, because that's much easier to test: Spin up and tear down a million pthreads in C, and see how long that takes and how much memory it takes. Then spin up and tear down a million processes in C and see your computer grind to a halt until you kill the process that is starting the processes, if you can even get your computer to do that without power-cycling. It's <50 lines of code for each, so I'm eagerly waiting for your response! Notably, my confidence here comes from the fact that I don't generally get into performance arguments without having actually tested what I'm saying. I've written this code before--it's what I do whenever I'm checking out a new programming language or threading library. Given the complexity of modern computers, nobody really can predict how a program will behave without testing it (except maybe in assembly) there's just too many variables. So you should stop doing that. If you decide to try the same thing in Java (the other language mentioned), probably drop the number of threads/processes down to 100,000, since Java's lightweight threads aren't quite as efficient. 100,000 processes will probably still be enough to crash your computer. I'm sure you can find some language/library which implements threads particularly inefficiently, so let's stick to pthreads/C and avoid that straw man. EDIT: Here ya go, I had ChatGPT write this one for ya: #include <stdio.h> #include <pthread.h> #include <unistd.h> void* threadFunction(void* arg) { // Sleep for 10 seconds sleep(10); pthread_exit(NULL); } int main() { int numThreads = 1000000; pthread_t threads[numThreads]; // Create threads for (int i = 0; i < numThreads; i++) { int result = pthread_create(&threads[i], NULL, threadFunction, NULL); if (result != 0) { printf("Failed to create thread %d\n", i); return 1; } } // Join threads for (int i = 0; i < numThreads; i++) { int result = pthread_join(threads[i], NULL); if (result != 0) { printf("Failed to join thread %d\n", i); return 1; } } return 0; } And... #include <stdio.h> #include <sys/types.h> #include <sys/wait.h> #include <unistd.h> int main() { int numProcesses = 1000000; pid_t childPID; // Create processes for (int i = 0; i < numProcesses; i++) { childPID = fork(); if (childPID < 0) { printf("Failed to create process %d\n", i); return 1; } else if (childPID == 0) { // Child process sleep(10); return 0; } } // Wait for all child processes to finish int status; pid_t pid; while ((pid = wait(&status)) > 0); return 0; } It looks like the latter just crashes the program without taking down my whole machine now, which is an improvement over the last time I tried this with processes.
- kzrdude 3y agoAnother new feature for 3.12 is per-subinterpreter GIL which is a middle ground. It would offer isolated interpreter threads without gil.
- usrbinbash 3y ago> you're already getting free multi-core. Please explain: In what sense is the overhead of starting actual OS processes, and relying on IPC "free", compared to running threads or even greenlets, and using shared process memory?
- csmpltn 3y ago> In what sense is the overhead of starting actual OS processes, and relying on IPC "free" With Python's current multiprocessing utilities - you get a big discount by not having to write thread-safe code, or worry about synchronization, despite the GIL still being there. Very broadly speaking, it's "free" in the sense that the OS handles parallelism automatically at the process-level, and provides a simple communication mechanism between those processes through standard APIs. It also reduces potential attack surfaces (although this is a lesser argument). It's also "free" in the sense that you don't need to re-write large parts of the VM, as-well as all supported libraries, and teach the entire Python community how to safely write and test multi-threaded code (something I bet upwards of 75% of the people using Python today won't manage) to support this specific form of parallelism. If your goal is to run code (whether IO bound or CPU bound) in parallel - Python has the means to do that already today, without removing the GIL.
- usrbinbash 3y ago> It's "free" in the sense that you don't need to re-write large parts of the VM In that sense, never updating python again is "free" as well, because it would save the python devs the trouble of changing the interpreter. And yet I think we can all agree that Python benefits from the fact that we no longer use Python 3.5 > and teach the entire Python community how to safely write multi-threaded code People who don't write threaded code don't need to worry about it. And people who write threaded code in python already need to worry about writing thread-safe code. The GIL doesn't prevent race conditions between individual python instructions. > Python has the means to do that already And as outlined above, these means are no suitable replacement for true thread based parallelism.
- 3y ago
- soulbadguy 3y ago> If python is being used as a scripting language And when it's not ?
- commonlisp94 3y agoWhen you write the entire system in python instead of using python to coordinate programs and libraries written in other languages.