4 ms·
I have a few questions about computer architecture for you hardware folks. These many core architectures get so complex. I'm a lowly scientist that does some fa
by earthscienceman 3y ago
I have a few questions about computer architecture for you hardware folks. These many core architectures get so complex. I'm a lowly scientist that does some fairly heavy compute in both interpreted languages (python) and compiled languages (c/julia). Two questions:
When these 'base clocks' are listed, would a lower core CPU with a higher base clock run more calculations at its max thread count than a higher core cpu? Say 64 threads on a 32 vore 4Ghz CPU vs a 2.5Ghz 96 core monster. For my use case I'm often running a serial computation on a fixed dataset that can be threaded per data point. So if I have 32 data points it's only that parallel. Something between embarrassingly parallel and completely serial.
And second. For these massive caches, if I'm doing a massive serial calculation does an individual core get to use the full cache if it's being thrashed? As in, are there cache benefits to running a computation on these huge CPUs vs some overclocked 8 core beast?
- bick_nyers 3y agoAs others have mentioned, it's complicated. This site will be your best friend here: https://openbenchmarking.org/tests https://openbenchmarking.org/tests
- Iulioh 3y agoWellll....it depends ------- When these 'base clocks' are listed, would a lower core CPU with a higher base clock run more calculations at its max thread count than a higher core cpu? Say 64 threads on a 32 vore 4Ghz CPU vs a 2.5Ghz 96 core monster ----- You know that GPUs are just CPUs with a massive core count, right? Like in the 1000s of cores So depends on what calculations you need done, there is no answer. It depends on how the program uses the cores, not even only what calculations you need done
- effie 3y agoGPU cores are not CPUs, otherwise we could run Linux and a shell on them. GPU core is much simpler than a CPU, that's why there can be so many of them on the die.
- jfindley 3y agoClock speed isn't a particularly meaningful measurement anymore, and hasn't been for years. For example, an AMD Genoa chip, depending on SKU, may have fairly comparable base/boost clock speeds compared to an Intel Sapphire Rapids - but in practice the single-core performance of the Intel is going to be substantially better for most code. Your cache question doesn't really have a simple answer either. E.g. an AMD CPU is split into different CCXs. To simplify somewhat, each core is broken up into several smaller compute units, with their own caches and memory controller. Intel has a completely different ring-based approach that's harder to summarise in once sentence. Overall though, for the sort of work you're describing the limiting factor is often memory bandwidth, not raw compute. Different platforms have very different membw/core figures, and I suspect if you started measuring that then you'd find it easier to predict your codes performance.
- adrian_b 3y agoWhile at equal clock frequency the Intel CPUs are a little faster in single-thread applications, their main weakness is that at equal power consumption their clock frequencies are much lower in multi-threaded applications, which leads to much lower multi-threaded performance. This can be easily noticed when comparing the base clock frequencies, which are more or less proportional with the actual clock frequencies that will be reached in multi-threaded applications. For instance a 7950X has 4.5 GHz versus the 3.2 GHz of 14900K. Similar differences are between Epyc and Xeon and between Threadripper and Xeon W. In desktop CPUs Intel can hide their very poor multi-threaded performance by allowing a much higher power consumption. However this method does not work for server and workstation CPUs, because these already have the highest TDP that is possible with the current cooling solutions, so in servers and workstations the bad Intel MT performance is much more visible. Intel hopes that this will change in 2024, when they will launch server and workstation CPUs made with the new Intel 3 CMOS process. In the absence of actual benchmarks, a good proxy for the multi-threaded performance of a CPU is the product between the base clock frequency and the number of cores. For Intel hybrid CPUs, an E-core should be counted as 0.6 cores. For example a Threadripper 7960X should be expected to be (24 cores x 4.2 GHz) / (16 cores x 4.5 GHz) = 1.4 times faster than a 7950X in multi-threaded applications that are limited by the CPU cores (but twice faster in applications that are limited by the memory throughput).
- icegreentea2 3y ago> And second. For these massive caches, if I'm doing a massive serial calculation does an individual core get to use the full cache if it's being thrashed? As in, are there cache benefits to running a computation on these huge CPUs vs some overclocked 8 core beast? On AMD chip's (I just dunno about the Intel architecture), each chip is divided up into chiplets with their own set of cores and L3 cache. For these Threadrippers, those will be 8 physical cores, and 32MB of L3 cache. Each core can access the L3 cache within their own chiplet only. You'll need to dig into the enabled cores+cache arrangements for any particular chip you might be interested in to figure out what will be good for your workloads.
- kllrnohj 3y ago> When these 'base clocks' are listed, would a lower core CPU with a higher base clock run more calculations at its max thread count than a higher core cpu? Say 64 threads on a 32 vore 4Ghz CPU vs a 2.5Ghz 96 core monster. The base clock is only when all cores are loaded, and especially on Threadripper / Ryzen is far from meaningful in practice as they will permanently turbo. So it's very possible that the 96 core threadripper and the 32 core threadripper when asked to run the same 32 threads will actually end up running around the same clockspeed. See for example this chart on the 5950X from anandtech: https://images.anandtech.com/doci/16214/CoreFreqScale-5950v3950-2a.png https://images.anandtech.com/doci/16214/CoreFreqScale-5950v3... Note that there's not 2 frequencies, there's a whole range of them and that range varies by the actual workload demands.
- sakras 3y ago> serial computation on a fixed dataset that can be threaded per data point. So if I have 32 data points it's only that parallel The general rule of thumb is that within a generation, higher clock speed yields better performance per core. If you only have 32 data points, then you will probably get better performance with the 4 GHz 32-core CPU. > if I'm doing a massive serial calculation does an individual core get to use the full cache Your single core will get to use the entire L3 cache, but L2 and L1 caches are per-core and so your single core doing the work will not have access to those. So yes, there conceivably could be a benefit due to the larger L3 cache. On a broader note, these kinds of factors (frequency, cache size, parallelism) tend to be extremely workload-specific and unpredictable, so the only real way to find out what's faster is to measure your specific workload.