13 ms·
How many CPU cores can you use in parallel?
- BerislavLopac 3y ago> Last updated 15 Dec 2023, originally created 18 Dec 2023 I respect Itamar even more after realising that he has invented the time machine... :o
- not_your_vase 3y agoThis happens when you can't sync your tasks correctly in a multi-threaded environment. Synchronous tasks happen in wrong order.
- pixelpoet 3y agoKnock, knock. Race condition! Who's there?
- BerislavLopac 3y agoThere are only two hard problems in distributed systems: 2. Exactly-once delivery 1. Guaranteed order of messages 2. Exactly-once delivery [0] [0] https://twitter.com/mathiasverraes/status/632260618599403520 https://twitter.com/mathiasverraes/status/632260618599403520
- deleted 3y ago[deleted]
- nolongerthere 3y agolol that’s a Wordpress bug, I don’t remember how I hit it, but I’ve done it before too. Might have to do with not having the datetime set correctly on the server.
- dale_glass 3y ago> But there is something unexpected too: the optimal number of threads is different for each function. Nothing unexpected there. Amdahl's Law in its glory. A fast running function finishes fast, and coordinating the job's execution over many cores requires doing work every time something finishes. If you split your job into chunks in the microsecond time range, then you'll be handling lots and lots of tiny minutia. You want to set up your tasks such that each thread has a good amount of stuff to chew through before it needs to communicate with anything.
- itamarst 3y agoThat's not it. I updated the article with an experiment of processing 5 items at a time. The fast function doing 5 images at a time is slower than the slow function doing 1 image at a time (24*5 > 90). If your theory was correct, we would expect the optimal number of threads for the fast function processing 5 images at a time to be similar to that of the slow function processing 1 image at a time. In fact, the optimal threads in this case (5 images at a time) was 20 for slow function, 10 for fast function, so essentially the same as the original setup.
- rewmie 3y ago[dead]
- bogwog 3y agoIs it a caching thing? The slow version seems less cache efficient, so if it is waiting due to cache misses, that could create an opportunity for something else to get scheduled in.
- gmm1990 3y agoI doubt it the slow version uses division instead of bit shifting. My guess would be the fast version saturated like i/o or some non cpu portion of the processor and the division one was bottle necked by the division logic in the processor.
- bogwog 3y agoBut it's iterating through the result vectors twice, so that's basically guaranteed to miss. Moving the threshold check into the loop above would at least eliminate that factor. Maybe division vs bit shifting does play a factor, but it's hard to compare that while the cache behavior is so different.
- cogman10 3y agoRecalibrate how you feel about division and multiplication. It turns out, integer division on new processors is a 1 cycle process (and has been for a while now). Most of the multicycle instructions now-a-days are things like SIMD and encryption.
- mihaic 3y agoStarting from Python 3.13 there should be a new method os.process_cpu_count() that aims to get the actually available number of cores our process can run on.
- itamarst 3y agoNeat! Unfortunately at the moment it still ignores cgroups, it's just a wrapper around sched_getaffinity(). https://github.com/python/cpython/blob/6a69b80d1b1f3987fcec3300c5dc879c6e965079/Lib/os.py#L1143 https://github.com/python/cpython/blob/6a69b80d1b1f3987fcec3...
- mihaic 3y agoTrue, I'm wondering if they might change the implementation for this actually, since the function name is pretty agnostic.
- onetimeuse92304 3y agoPersonally, I dislike configuration parameters that let future admins of the system change parameters like concurrency of certain processes, etc. A lot of the time even I can't tell what is going to be the optimum setting. So for a number of years, rather than expose a configuration parameter, I am building in a performance test that establishes the values of these parameters. This is a little bit of search and a hill climb through the multidimensional space of all relevant parameters. It is not perfect, but in my experience it is almost always better than I can do manually and always better than an operator without intimate understanding of the system. The results are cached in a database and can be re-executed if change to configuration is found (just take a number of parameters from your environment, version of the application, etc. and hash the value and check if you already have parameters in the database).
- eropple 3y agoThis is interesting - how long does it take to write? Is it reusable? I'd be really interested in a blog post about this.
- onetimeuse92304 3y agoI am not big about writing blog posts. But now that I think about it, I could write a small library to automate most of it while delegating some tasks to the application (providing functionality to be tested, parameters and possible ranges to be searched, a way to save/restore parameters to/from the database, provide fitness function, etc.) It is essentially a small benchmarking framework with some added functionality. I think it could be quite useful to some people.
- arp242 3y agoThis works well for expensive long-running things, but works less well for more short-lived programs, and even for expensive long-running things there may be more considerations than "what's the fastest value?" I intentionally set the default jobs of "make" to "4" rather than "8" because I don't want to utilize my full system when compiling stuff, so I can still use it for other things at a reasonable speed. There's lots of reasons you might want to do this: on a web platform you don't really want to use up all resources by the one expensive thing, but what's "reasonable" and "desired" often differs – often there is no objectively best "optimum setting".
- mrlonglong 3y agoMore cores are excellent for building large projects. My threadripper is worth every penny in software development.
- npoc 3y agoThe 7950X eats large C++ projects for breakfast, especially those with lots of boost includes. Every logical core counts.
- Skunkleton 3y agoI'm still waiting for multithreaded c++ compiler. The project I work on has some nasty templating, and there are a few compilation units that take a long long time.
- mrlonglong 3y agoWouldn't refactoring to reduce the templating be helpful? Templates are an horror if it's out of control. I have a small project that uses templates and there is a bug with one of the templates in it that I haven't been able to resolve yet. I do look at it occasionally see if I get anywhere with it. Frankly I found Rust a lot easier to work with.
- Skunkleton 3y agoI will probably get a multithreaded compiler before I can convince people to stop abusing templates shrug
- bogwog 3y agoAdd mold into the mix, and you're in C++ heaven. https://github.com/rui314/mold https://github.com/rui314/mold
- mrlonglong 3y agoI'll give that a try. One of the big ones is Chromium, and it still takes 2½ hours to build with all 48 cores at full throttle.
- deleted 3y ago[deleted]
- Kon-Peki 3y ago> you can spend a little bit of time at the beginning to empirically measure the optimal number of threads, perhaps with some heuristics to compensate for noise. I'd like to point out that there is a lot of stuff published on ArXiv; this appears to be a very active research subject. Don't start from scratch :)
- dahart 3y ago> our faster function could take advantage of no more than 8 cores; beyond that it started slowing down. Perhaps it started hitting some bottleneck other than computation, like memory bandwidth. @itamarst Yes, this in interesting, you should profile it and get to the bottom of the issue! It seems like in my experience that being limited by hyperthreading or instruction-level parallism is relatively rare, and much more often it’s cache or memory access patterns or implicit synchronization or contention for a hardware resource. There’s a good chance you’ll learn something useful by figuring it out. Maybe it’s memory bus contention, maybe it’s cache, maybe numba compiled in something you aren’t expecting. Worth nothing that using 20 on the fast test isn’t that much slower than using 8. A good first guess/proxy for number of threads to use is the number of cores, and that pays off in this case compared to using too few cores. Out of curiosity, do you know if your images are stored row-major or column major? I see the outer loop over shape[0] and inner loop over shape[1]. Is the compiled code stepping in memory by 1 pixel at time, or by a whole column? If your stride is a column, you may be thrashing the cache. I’d also be curious to hear how the speed of this compiled code compares to a numpy or PIL image threshold operation, if you happen to know.
- itamarst 3y agoNumPy default is that you iterate over the earlier dimensions first. The slow code is likely at least partially slow due to branch misprediction (this is specific to my CPU, not true on CPUs with AVX-512), see https://pythonspeed.com/articles/speeding-up-numba/ https://pythonspeed.com/articles/speeding-up-numba/ where I use `perf stat` to get branch misprediction numbers on similar code. With SIMD disabled there's also a clear difference in IPC, I believe. The bigger picture though is that the goal of this article is not to demonstrate speeding up code, it's to ask about level of parallelism given unchanging code. Obviously all things being equal you'll do better if you can make your code faster, but code does get deployed, and when it's deployed you need to choose parallelism levels, regardless of how good the code is.
- dahart 3y ago> when it's deployed you need to choose parallelism levels, regardless of how good the code is. Yes, absolutely, exactly. That’s why it can be really helpful to pinpoint that cause of slowdown, right? It might not matter at deployment time if you have an automated shmoo that calculates the optimal thread load, but knowing the actual causes might be critical to the process, and/or really help if you don’t do it automatically. (For one, it’s possible the conditions for optimal thread load could change over the course of a run.)
- omgtehlion 3y ago> Intel i7-12700K processor Wait! You can’t benchmark on a CPU with unpredictable speed and mix of slow and fast cores. All kinds of effects come into play here, and the code itself is not the most prominent among them. To measure which _code_ is better, you should use real SMP machine with fixed clock speed and turned off HT. On the machine from TFA you are just fighting with Intel’s thermal smarts and OS scheduling shenanigans. (edit: you can use the same machine, but configure it in BIOS. I, myself, use i9-12900k fixed to 5ghz@8P-cores as "Intel testing machine")
- sspiff 3y agoExcept then you're not testing on the hardware configurations people will actually run the software on. People do run software with HyperThreading on and on E cores and with Turbo boost and throttling. Having code that behaves better in these dynamic, heterogeneous environments is a net benefit to the user.
- omgtehlion 3y agoIt is easier to deal with all the issues separately. First your algo (in single thread), then threading, then memory accesses and cache effects. After you sort everything in your control, you try to deal with the real world (like counting only “real” cores, ignoring/or enjoying hyperthreads, OS scheduling, but most of these are quite unpredictable, and if you get your code run faster in 12700k, the same setting will be slower on, say, 5800X or on server machines, while basic stuff speeds up the program on any hardware).
- spenczar5 3y agoThose issues are not separable. The performance of memory and caches can determine which algorithm is best. This is the basis of the entire field of cache-aware algorithms, which routinely beat the pants off of theoretically superior algorithms. In my experience (which is in video transcoding and in research astrophysics - both domains where it matters!), if you really need to squeeze out performance, you have to design with the target platform available for profiling and benchmarking from the beginning. Edit to add: I agree wholeheartedly with your top-level comment! I just am perhaps more extreme than you; I don’t think “laptop benchmarks” can ever be twisted into being useful.
- kristjansson 3y ago> (you can also use perfplot, but note it’s GPL-licensed) Surly using a GPL-licensed development tool cannot affect the licensing of the project it's being used to develop?
- deleted 3y ago[deleted]
- mroche 3y agohttps://www.gnu.org/licenses/gpl-faq.html#GPLOutput https://www.gnu.org/licenses/gpl-faq.html#GPLOutput Is there some way that I can GPL the output people get from use of my program? For example, if my program is used to develop hardware designs, can I require that these designs must be free? In general this is legally impossible; copyright law does not give you any say in the use of the output people make from their data using your program. If the user uses your program to enter or convert her own data, the copyright on the output belongs to her, not you. More generally, when a program translates its input into some other form, the copyright status of the output inherits that of the input it was generated from. So the only way you have a say in the use of the output is if substantial parts of the output are copied (more or less) from text in your program. For instance, part of the output of Bison (see above) would be covered by the GNU GPL, if we had not made an exception in this specific case. You could artificially make a program copy certain text into its output even if there is no technical reason to do so. But if that copied text serves no practical purpose, the user could simply delete that text from the output and use only the rest. Then he would not have to obey the conditions on redistribution of the copied text. --- https://www.gnu.org/licenses/gpl-faq.html#WhatCaseIsOutputGPL https://www.gnu.org/licenses/gpl-faq.html#WhatCaseIsOutputGP... In what cases is the output of a GPL program covered by the GPL too? The output of a program is not, in general, covered by the copyright on the code of the program. So the license of the code of the program does not apply to the output, whether you pipe it into a file, make a screenshot, screencast, or video. The exception would be when the program displays a full screen of text and/or art that comes from the program. Then the copyright on that text and/or art covers the output. Programs that output audio, such as video games, would also fit into this exception. If the art/music is under the GPL, then the GPL applies when you copy it no matter how you copy it. However, fair use may still apply. Keep in mind that some programs, particularly video games, can have artwork/audio that is licensed separately from the underlying GPLed game. In such cases, the license on the artwork/audio would dictate the terms under which video/streaming may occur.
- fizzynut 3y agoI think the author has discovered memory bandwidth. When you have a simple function and just scale the number of cores it's easy to hit.
- jjslocum3 3y agoI understand this is a Python-centric source, but without having done my homework I'd have thought Python wouldn't be a particularly great language for dealing with these low level concerns. Wouldn't it be much easier in C? In java it's as simple as Runtime.getRuntime().availableProcessors()
- mort96 3y agoI mean other than being wrapped in a useless singleton, that's the same as the 'os.cpu_count()' mentioned in the beginning of the article.
- navels 3y agoRelated: fascinating deep dive of the Python GIL by David Beazley from PyCon2010: https://www.youtube.com/watch?v=Obt-vMVdM8s https://www.youtube.com/watch?v=Obt-vMVdM8s
- shanemhansen 3y agoThe problem of "how many CPUs should I use?" is really only answerable empirically in the general case. The problem of "how many CPUs are available?" is a little more tractable. Currently when running podman, the cpus allocated seems to be available in the /sys fs. I wonder if it's the same under k8s? podman run --cpus=3.5 -it docker.io/library/alpine:latest /bin/sh / # cat /sys/fs/cgroup/cpu.max 350000 100000
- o11c 3y agoPer cgroups v2 documentation [0], `cpu.max` exists for all but the root cgroup. If it doesn't exist, that means the kernel at least is putting no restrictions compared to what the affinity says (use `nproc` to call `sched_getaffinity` from the command line). I'm not sure what hypervisor-based throttling exposes. *** That said, there are several rules of thumb to guess the optimal number of CPUs, even before you measure - and to figure out how to improve the situation (since profiling is even harder for threaded code): * if you are contending for a shared resource, even relatively rarely [1] (that is, assuming you're already using per-thread resources when reasonable and only sharing big chunks), you're likely to hit a limit for the maximum number of CPUs regardless. Numbers I've seen are often 20 or lower. Using a "nested" approach for resource sharing can improve this, but this may require pinning threads (which in turn threatens starvation - REMEMBER, YOU ARE NOT THE ONLY PROGRAM IN THE WORLD!). * if you are truly CPU-bound, microoptimizing to avoid (mostly: random-access) memory stalls (usually only possible for purely numerical code), logical CPUs are useless so you should limit yourself to the physical CPU count. Usually you should have an idea if this is the case. * if your program is actively using an amount of memory similar to a shared cache size, using fewer CPUs is often better. This is not restricted to the exact counts of logical and physical CPUs, except for caches that are physical-CPU-specific (L1 almost always is, and L2 usually is these days for P-cores but not E-cores. L3 - and, for that matter, total RAM size to avoid swapping - is the one where the optimal number of CPUs varies arbitrarily in practice) * if your code does even a small number of unpredicted loads (typical object-oriented code does far more than that!) logical CPUs are easily a win. Most code belongs here so you can assume even if you're not familiar with it, unless one of the hard-to-know points applies. [0]: https://www.kernel.org/doc/html/latest/admin-guide/cgroup-v2.html https://www.kernel.org/doc/html/latest/admin-guide/cgroup-v2... [1]: https://en.wikipedia.org/wiki/Amdahl's_law https://en.wikipedia.org/wiki/Amdahl's_law
- dr_kiszonka 3y agoPythonSpeed is one of my favorite websites. However, this article leaves me with more questions than answers. (I do appreciate the benchmarking code.) For example, at one point, the author mentions that hyper-threading can be disabled in BIOS. Should I disable it? Based on the author's own description, it sounds like hyper-threading is pretty useful.
- owlbite 3y agoLike much of the advice "it depends" and "measure it". Basically hyper-threading shares a whole load of resources between threads. If your code ends up contending on those resources you probably run slower. If your code is blocked due to lack of instruction-level parallelism, it probably runs faster. Generally an expert-optimized numerical code (e.g. machine learning) is going to fall into the first camp as it's pretty easy to kill the ILP bottlenecks in that sort of code.
- duohedron 3y agoWhen I wrote some parallel code in C, enabling hyper-threading resulted only 30% increase in speed, which not too little, but you would expect more for doubling the thread count. I don't remember what was the bottleneck, but it was not I/O. On an other occasion a consultant for a HPC consultant advised us to disable some physical cores for the optimal performance of a geological simulator application, because the core/cache ratio is higher that way.
- gpderetta 3y ago30% increase in performance from hyperthreading is impressive.
- IshKebab 3y agoNo. It's a performance optimisation and generally makes things faster. The only reason to disable it is because of Spectre security stuff.
- CyberDildonics 3y ago
- gnufx 3y agoUse hwloc[1] to determine the hardware topology of the system, and pin processes/threads appropriately. For a heterogeneous system, you presumably want to avoid low-performance cores. OpenMP has a standard way of assigning threads to cores. Otherwise use hwloc, or perhaps likwid[2] explicitly. Most things turn out to be memory-limited, and getting the algorithm right may win more than simple parallelism; GEMM is a case in point. As ever, profile first, and worry about thread numbers after you've pinned appropriately and looked for simple things like initialization interacting with memory policy. The main HPC performance tools support Python and native code; Threadspotter was good for threading/cache analysis, but seems moribund. Note that SMT on POWER, for instance, is different from x86. (This is general HPC performance engineering; I haven't considered the code in point.) 1. https://www.open-mpi.org/projects/hwloc/ https://www.open-mpi.org/projects/hwloc/ 2. https://github.com/rrze-likwid/likwid https://github.com/rrze-likwid/likwid
- 1letterunixname 3y agoRemember kids, classical x86 hyperthreading goes freeganning for additional under-scheduled execution units of each physical core to create another virtual core that's not going to be as fast as doubling the count of physical cores.
- EVa5I7bHFq9mnYK 3y agoIn my programs if a physical core performs 1000 operations/sec, with an added HT core it performs 1070 ops/ second, not 2000 as I naively expected. I understand you should specifically program so that the main core performs integer arithmetics, and the HT core - floating point arithmetics, to get closer to 2000.
- godelski 3y agoOne of the hard things about writing parallelized programs is that you may want different algorithms for different environments. The other day I joked about CUDA being harder than rocket science[0] and this is part of it. The thing with high parallelism is that you can't just think about your program in terms of clock and total memory. You have to consider all parts of the system, and this is likely the large reason most programs don't have high parallelism despite computers being many cores for decades. OpenMP is going to have different settings than OpenMPI. How you share the memory is an essential part of the algorithm you're going to write. There's also issues about real cores and hyperthreads/logical. I find that in most computation I never want to use the number of logical cores but only rely on physical cores. You can also get a double descent like behavior[1]. The author here seems to be running into an even newer problem that is the difference between performance cores and efficiency cores. I'm not actually sure that this is surprising and makes the headline feel clickbaity. For more, this is actually part of why CUDA is making a big boom lately. One of my first forays into CUDA was trying to write some kernels to speed up GEANT4 simulations, which are not too dissimilar from path tracing but they are more computationally heavy. So extremely parallelization but the operations per cycle wasn't high enough back then so it was still better to use CPU (I'm sure there were other bottlenecks as well but did find someone else confirming my results). This goes for lots of scientific compute too, like a lot of mesh solvers and FEA (finite element analysis: how you simulate physical parts). But not IPC has gone way up (AND cores increase!! AND bandwith! AND memory!) so now we start seeing these things start to leverage graphics cards a lot more. In HPC and it is worth noting that the big bottleneck actually isn't compute, it's throughput. More data can be created than can be processed. Many teams are switching to "in situ" visualization/analysis methods where you offload your data to another machine which does that processing. You also simulate at FAR higher resolution than you visualize or even analyze. So if you're interested in tackling problems that are becoming more and more important, this is one of them and involves both hardware and software. [0] https://news.ycombinator.com/item?id=38678066 https://news.ycombinator.com/item?id=38678066 [1] Say you have 16 physical cores with hyperthreading so 32 logical cores. Your program speeds up as number of processes -> 16. But when you go to 17 cores you have a big jump (loss in speedup) and follow a shallower curve as number of processes -> 32. Edit: If you're not on an HPC machine, I actually might suggest setting your max parallelism to (number_of_physical_cores - 1) because speed difference often is trivial (lots of asterisks here) and you have a core available to kill the process. More noob you are, the stronger this recommendation.
- deleted 3y ago[deleted]