6 ms·
Talking of the new chips, it seems interesting to me to note the difference in cache sizes. I know it is not a 1:1 comparison because of architecture difference
by actuator 5y ago
Talking of the new chips, it seems interesting to me to note the difference in cache sizes. I know it is not a 1:1 comparison because of architecture difference but the 16 core 5955WX has
L1: 512 KB, L2: 8 MB, L3: 64 MB
compared to M1 Ultra's from the Geekbench data
L1: 128 KB instruction, 64 KB data, L2: 4 MB
I have mostly forgotten my microprocessor architecture lectures but it seems interesting that even after being able to cache more data near a core, AMD is not gaining much. Maybe packing too much cache increases latency of access or the gains simply go away beyond a certain size.
Edit:
Maybe even it is coming down to the cache layout. Does anyone know if the cache fetch times for the named levels are roughly the same across architectures?
For x86, I remember L1 being a single cycle fetch and L2 being 10-20x slower than L1
- PragmaticPulp 5y agoThe M1 Ultra has a significant memory bandwidth and latency advantage due to the type of memory used and the way the memory is connected. They use the same memory for the CPU and GPU, so the memory interface was optimized more like a GPU and the CPU benefits in a few memory-constrained benchmarks (machine learning, AES-XT streaming) The flip side is that you're limited to 128GB combined memory for the CPU and GPU on the M1 Ultra, whereas a comparable Threadripper Pro system will take up to 16 times as much (2TB) and you can upgrade it whenever you feel like. It will be interesting to see how much RAM Apple offers on the upcoming M1 Mac Pro parts.
- actuator 5y ago> significant memory bandwidth and latency advantage This is to cache or the main memory? I remember x86 based PCs taking 100ns for RAM access. Is it faster in ARM? > They use the same memory for the CPU and GPU, so the memory interface was optimized more like a GPU and the CPU benefits in a few memory-constrained benchmarks (machine learning, AES-XT streaming) This would mean each core having some dedicated RAM section better connected. Wouldn't this be more like L3 section in x86 but bigger? Maybe this is the advantage of having everything on the same die. Beyond a certain size gains should also flat out, no?
- humanwhosits 5y ago> Is it faster in ARM? Not in the general case, but specifically with how Apple has brought the ram chips so physically close to the cpu cores
- actuator 5y agoI wonder what stops AMD, Intel in trying this, except loss of modularity. It is not like the architecture will change by bringing it into the die. As far as I remember, CPU is the only thing that does RAM access even in x86
- kube-system 5y agoI would be surprised if they’re not evaluating this. The only difference I can see is that it complicates the number of SKUs they’d need to provide to their customers. It’s a little easier for Apple in this regard because they’re also building the final machines, whereas Intel/AMD are making chips that are going in a wider range of devices.
- wmf 5y agoThe physical distance confers no latency advantage.
- actuator 5y agoIs this because of signal speed anyway reaching almost speed of light?
- willis936 5y agoMaybe 70% the speed of light. That's a handful of cycles at a few GHz for any SDRAM channel. The latency limitation is caked in to the scanning, strobed nature of SDRAM. It's dense and cheap, but it will never respond faster than ~20 ns (100+ cycles).
- brigade 5y agoBandwidth yes, latency no. Wiring out 32 channels of DDR5 to 16 slots might not be feasible, but latency-wise Anandtech's measurements suggest the M1 Max actually has a bit higher latency to memory than e.g. Icelake-SP
- wmf 5y agoLooks like you're comparing per-chip numbers vs. per-core numbers.
- actuator 5y agoGeekbench mentioned it under just the processor not core. 4 MB of per core L2 sounds too much tbh
- kinghajj 5y agoI'm almost certain that the L1 caches on the M1 ultra are being reported incorrectly. That 512 KiB on the AMD CPU is the sum of data and instruction cache across all cores. M1 performance cores have 192 KiB of instruction cache and 128 KiB of data cache per core; and the efficiency cores have 128/64 KiB of instruction and data caches, respectively. 16*(192+128)+4*(128+64) = 5888 KiB of L1 cache on the M1 ultra.
- actuator 5y agoWow, if that's true then that is such a massive advantage. L1 is a single cycle fetch cache if it is like x86. So, individual cores can do compute so much better by fitting more data at once.
- wmf 5y agoL1 hasn't been a single cycle for a long time, like decades.
- actuator 5y agoHmm, thanks; that seems interesting. I guess I need to read up more. On searching in Google someone[1] was quoting Xeon L1 fetch as approximately 4 cycles. I don't know if this is an average across branch prediction hit/miss, I will look for source of these numbers and try to read what has changed. [1] https://stackoverflow.com/a/4087331 https://stackoverflow.com/a/4087331
- moonchild 5y agoYes, 4-5 cycles is to be expected. Also, larger caches generally trade off latency, so I would not be surprised if apple's chips were slower still.
- 95014_refugee 5y agoThe cycle count is largely irrelevant in a speculative, pipelined core, as long as it can speculate far enough ahead that the delay is masked by other activity. 4-5 cycles is nothing in the scheme of things.
- sharikous 5y agoIn addition to what the other comments pinpoint there is a SLC (system level cache) in the L1 that does the job of the L3 cache
- GeekyBear 5y ago> compared to M1 Ultra's L1: 128 KB instruction, 64 KB data, L2: 4 MB That doesn't quite match up with what we know about the two M1 Max chips that make up the Ultra. Here are the resources per M1 Max, so multiply by two. >On the core and L2 side of things, there haven’t been any changes and we consequently don’t see much alterations in terms of the results – it’s still a 3.2GHz peak core with 128KB of L1D at 3 cycles load-load latencies, [192k L1 instruction cache], and a 12MB L2 cache. Where things are quite different is when we enter the system cache, instead of 8MB, on the M1 Max it’s now 48MB large https://www.anandtech.com/show/17024/apple-m1-max-performance-review/2 https://www.anandtech.com/show/17024/apple-m1-max-performanc...
- xxs 5y ago>L1: 128 KB instruction, 64 KB data, L2: 4 MB Writing this would make no sense, knowing how many (more) transistors M1 has. In very simple terms one bit of SRAM is 6 transistors (+some extra for addressing). Also comparing M1 Ultra to consumer grade chips is quite pointless, x86-64 consumer grade ones have just 2 memory channels.
- hajile 5y agoM1 Max P-cores have 128 L1 D-cache, 196kb I-cache, and 24mb L2. E-cores have 64kb L1 D-cache, 128kb L1 I-cache, and 4mb L2. There is also a shared 48mb system-level cache that serves as L3. AMD Zen 3 has 32kb D-cache and 32kb I-cache per core. Here's the real chart M1 Ultra AMD 5950X AMD 5995WX L1-D 1.13mb 512kb 2mb L1-I 1.78mb 512kb 2mb L1-total 2.91mb 1mb 4mb L2-total 56mb 8mb 32mb L3/SLC-total 96mb 64mb 256mb Total L2+L3 152mb 72mb 288mb Total Cache 155mb 73mb 292mb