4 ms·
You focus on rust rather than generalizing... If you are IO bound, consider threads. This is almost the same as async / await. What was missing above, and th
by jmspring 3y ago
You focus on rust rather than generalizing...
If you are IO bound, consider threads. This is almost the same as async / await.
What was missing above, and the problem with how most compute education is these days, if you are compute bound you need to think about processes.
If you were dealing with python concurrent.futures, you would need to consider processpooexecutor vs. threadpoolexecutor.
Threadpoolexecutor gives you the same as the above.
With multiprocessor executor, you will have multiple processes executing independently but you have to copy a memory space. Which people don't consider. In python DS work - multiprocessor workloads need to determine memory space considerations.
It's kinda f'd up how JS doesn't have engineers think about their workloads.
- baq 3y agoBackend JS just spins up another container and/or lambda and if it's too slow and requires multiple CPUs in a single deployment, oh well, too bad.
- zelphirkalt 3y agoThat is of course a huge overhead, compared to how other languages solve the problem.
- fennecfoxy 3y agoBackend JS does whatever the hecky you _want_ it to do. The Cluster module has been around for a long time: https://nodejs.org/api/cluster.html https://nodejs.org/api/cluster.html Tbf a lot of the time you're running it in a container and you allocate 1 vcpu there, only downside is maybe a little extra memory overhead. And for most lambdas I think they're suited to being single threaded (imo).
- danbruc 3y ago[...] if you are compute bound you need to think about processes. How would that help? Running several processes instead of several threads will not speed anything up [1] and might actually slow you down because of additional inter-process communication overhead. [1] Unless we are talking about running processes across multiple machines to make use of additional processors.
- zelphirkalt 3y agoI think you need to clarify what you mean by "thread". For example they are different things when we compare Python and Java Threads. Or OS threads and green threads. I think the GP was relating to OS threads.
- danbruc 3y agoI was also referring to kernel threads. If we are talking about non-kernel threads, then sure, a given implementation might have limitations and there might be something to be gained by running several processes, but that would be a workaround for those limitations. But for kernel threads there will generally be no gain by spreading them across several processes.
- josephg 3y agoRight; a process is just a thread (or set of threads) and some associated resources - like file descriptors and virtual memory allocations. As I understand it, the scheduler doesn’t really care if you’re running 1000 processes with 1 thread each or 1 process with 1000 threads. But I suspect it’s faster to swap threads within a process than swap processes, because it avoids expensive TLB flushes. And of course, that way there’s no need for IPC. All things being equal, you should get more performance out of a single process with a lot of threads than a lot of individual processes.
- Galanwe 3y ago> a process is just a thread (or set of threads) and some associated resources - like file descriptors and virtual memory allocations Or rather the reverse, in Linux terminology. Only processes exist, some just happen to share the same virtual address space. > the scheduler doesn’t really care if you’re running 1000 processes with 1 thread each or 1 process with 1000 threads Not just the scheduler, the whole kernel really. The concept of thread vs process is mainly a userspace detail for Linux. We arbitrarily decided that the set of clone() parameters from fork() create a process, while the set of clone() parameters through pthread_create() create a thread. If you start tweaking the clone() parameters yourself, then the two become indistinguishable. > it’s faster to swap threads within a process than swap processes, because it avoids expensive TLB flushes Right, though this is more of a theorical concern than a practical one. If you are sensible to a marginal TLB flush, then you may as well "isolcpu" and set affinities to avoid any context switch at all. > that way there’s no need for IPC If you have your processes mmap a shared memory, you effectively share address space between processes just like threads share their address space. For most intent and purposes, really, I do find multiprocessing just better than multithreading. Both are pretty much indistinguishable, but separate processes give you the flexibility of being able to arbitrarily spawn new workers just like any other process, while with multithreading you need to bake in some form of pool manager and hope to get it right.
- initplus 3y agoI think you are coming at this from a particular Python mindset, driven by the limitations imposed on Python threading by the GIL. This is a peculiarity specific to Python rather than a general purpose concept about threads vs processes.
- Hendrikto 3y ago> If you are IO bound, consider threads. This is almost the same as async / await. Only in Python. > if you are compute bound you need to think about processes. Also only in Python.
- cryptonector 3y agoIf you're using threads then consider not using Python. Or, just consider not using Python.
- fjdhdhdhfhfhf 3y ago^ kid who only writes python condescending to the engineers actually solving hard problems LOL