3 ms·
Point of Order! It's not L4, it's a ram cache. Data from L1-L3 isn't stored there, only what you have written to or read from ram. Your working set of data won
by undersuit 3y ago
Point of Order! It's not L4, it's a ram cache. Data from L1-L3 isn't stored there, only what you have written to or read from ram.
Your working set of data won't spill out from the L3 to "the L4" when it grows too large.
- Aissen 3y agoI'm not sure I understand the difference. Are we both talking about the "HBM Caching Mode" on this slide: https://www.servethehome.com/intel-xeon-max-9480-deep-dive-intel-has-64gb-hbm2e-onboard-like-a-gpu-or-ai-accelerator/intel-xeon-max-summary-slide/ https://www.servethehome.com/intel-xeon-max-9480-deep-dive-i... ?
- undersuit 3y agoThe memory caching, AFAIK, exist on different sides of explicit load and store instructions than CPU caching. It introduces subtle issues. You can see on page 4 a number of cache friendly benchmarks show little benefit for the HBM caching: https://www.servethehome.com/wp-content/uploads/2023/09/Intel-Xeon-MAX-Sample-Workload-Basket-696x416.jpg https://www.servethehome.com/wp-content/uploads/2023/09/Inte... 64GB of HBM2e as ram is more performant than 128GB of DDR5 with HBM2e cache, and often the cached variant has no speedup compared to a standard Intel configuration. Also OpenFOAM loves cache: https://www.phoronix.com/benchmark/result/amd_ryzen_7_5800x3d_linux_gaming_/ac34195db3dd.svgz https://www.phoronix.com/benchmark/result/amd_ryzen_7_5800x3...
- Aissen 3y agoI'd be curious to know more about the subtle issues (which I don't doubt there might be!). IMHO those results don't contradict the what I said. Of course if the workload entirely fits in the 64GB HBM, there's no point in using it as a cache, just use it directly. But if you need to address more RAM(any big DB, fs, etc.), and you don't want to manage the tier manually, then the caching mode could shine.
- undersuit 3y agoMemory caching will accelerate all requests through the memory controller. The extra 64MB of L3 in my 5800X3D wouldn't accelerate DMA between my memory and my GPU, this HBM cache should accelerate PCI-e devices like GPUs and Storage. It's a benefit in 2 socket configurations when the memory the CPU needs is connected to another socket. The data will be cached on the other socket's HBM for the entire system without filling up that other socket's L1-L3 cache for data it hasn't requested yet. Another way to see it is the last time Intel messed around with L4 cache under the EDRAM section. https://www.anandtech.com/show/9582/intel-skylake-mobile-desktop-launch-architecture-analysis/5 https://www.anandtech.com/show/9582/intel-skylake-mobile-des... The entire GPU and EDRAM complex get moved out of the CPU caching scheme.
- Aissen 3y agoOooh, I didn't think of the impact on DMA or 2-socket configurations, it's quite interesting indeed. Thanks !