4 ms·
Having substantially more L1 and L2 cache per core but no L3 has to be a massive part of why the M1 performance is so good. I wonder if Intel/AMD have plans to
by OldTimeCoffee 5y ago
Having substantially more L1 and L2 cache per core but no L3 has to be a massive part of why the M1 performance is so good. I wonder if Intel/AMD have plans to increase the L1/L2 size on their next generations.
- amackera 5y agoCan you help me understand why removing L3 cache would speed things up? Genuinely curious! Increasing L1 and L2 make intuitive sense.
- jessermeyer 5y agoI think the idea is by removing L3 has allowed for an increase of both L1/2.
- lanna 5y agoI think he meant "despite not having L3"
- cogman10 5y agoRemoving L3 frees up transistors to be spent on L1/L2. On a modern processor the vast majority of transistors are spent on caches. Why this might help, ultimately, because the latency for getting something from L1 or L2 is a lot lower than the latency from L3 or main memory. That said, this could hurt multithreaded performance. L1/2 are used for 1 core in the system. L3 is shared by all the cores and a package. So if you have a bunch of threads working on the same set of data, having no L3 would mean doing more main memory fetches.
- vbezhenar 5y agoApple will invent L3 for workstation-level CPU.
- duskwuff 5y agoWild theory: for workstation-class systems, the 8/16 GB of on-package memory becomes "L3", and main memory can be expanded with standard DIMMs.
- skunkworker 5y agoIf you could get a Mac Pro with 32 to 48 firestorm + 4 icestorm cores with tiered memory caching and expandable to 2TB+ DDR4/DDR5 DIMMs. That would be an impressive machine for the small amount of wattage it would draw from the wall.
- hajile 5y agoAMD went from 64kb in Zen 1 down to 32kb in Zen 2/3. Bigger isn't always better. It only matters if the architecture can actually use the cache effectively. M1 has a massive reorder buffer, so it needs and can use more L1 cache. It's pretty much that simple.
- monocasa 5y agoIt's more complicated on x86 because of the 4k page size. The L1 is heavily complicated if it is larger than the number of cache ways times the page size, since the virtual->physical tlb lookup happens in parallel. 8 way * 4kb = 32kb. AppleARM runs with a 16kb page size. 8 way * 16kb = 128kb
- wmf 5y agoIntel/AMD have decades of experience trading off L1, L2, and L3 so it's unlikely that there's a magic design they've overlooked.
- nknealk 5y agoTurns out, AMD did the opposite: https://www.tomshardware.com/news/amd-shows-new-3d-v-cache-ryzen-chiplets-up-to-192mb-of-l3-cache-per-chip-15-gaming-improvement https://www.tomshardware.com/news/amd-shows-new-3d-v-cache-r...