4 ms·
I am curious how much of the M1 gains can be realized by just slapping 16 gigs of cache on your die. I know that's a gross oversimplification, and I'm 100% just
by errantspark 5y ago
I am curious how much of the M1 gains can be realized by just slapping 16 gigs of cache on your die. I know that's a gross oversimplification, and I'm 100% just going off intuition here but it seems to me that memory access efficiency is what carries the brunt of the M1's gains.
- ericye16 5y agoIf you slapped 16GB of cache on your die, you would have a humongous die and the most expensive processor in the world. It would be very fast though.
- GhettoComputers 5y agoIs 1GB interesting enough? https://gadgettendency.com/amds-new-processors-will-have-nearly-1gb-of-cache-milan-x-will-receive-additional-cash/ https://gadgettendency.com/amds-new-processors-will-have-nea...
- smolder 5y agoAt some point the physical distance to parts of cache would mean you'd have rapidly diminishing returns on adding more. For 16GB it'd need to be some kind of tiered thing with nearer cache segments being quicker. Maybe you could have it present itself as a giant unified cache... sort of like what they did in the new IBM mainframes.
- formerly_proven 5y agoThe M1 parts actually have way worse memory latency (i.e. random accesses) than both AMD and Intel. A finely tuned Intel system has around 3x lower memory latency than an M1P/X. All of this is SDRAM, the S stands for synchronous and means that the memory is driven according to fixed timings in relation to a fixed bus clock. All LPDDR4X-4266 parts with the same timings perform exactly the same, whether they are soldered to the interposer or are 5 cm away on the board.
- errantspark 5y agoInteresting, I didn't know that the memory latency was worse. The bandwidth is still much higher right? In your estimation how much of the perf gains of the M1 chip are due to the increased memory bandwidth vs other optimizations/advancements?
- formerly_proven 5y agoI don't think you can single out one aspect when comparing designs which are so different, but I'd say most of the CPU performance is down to the core. There you have the M1 core which is (iirc) 10 pipes wide, much wider than both AMD and Intel, and a much fatter frontend as well to feed the beast. I think that's the main advantage the M1 has. It also has a larger L1 cache (twice as big as x86). These two are one of the few areas where x86 has an innate disadvantage because fattening up the frontend is much more complex on x86 compared to ARM, and you can't have a VIPT L1 cache larger than 64K with 4K pages, which is a hard-to-change default on x86, while the M1 by default uses larger pages (32K or something like that).
- errantspark 5y ago"x86 has an innate disadvantage [...] fattening up the frontend" this is inherent largely because of ISA complexity and variable width instructions? "M1 core which is [...] much wider than both AMD and Intel" The core width you're referring to is the decoder width yeah? As another poster pointed out the M1 has a large reorder buffer as well. Combined with the ability to index a larger L1 cache what I'm getting here is that the M1 can be a lot better about scheduling instructions (and perhaps even running non-interfering instructions in parallel on a given core (is that a thing?)) because the frontend has more power and space to do so. I guess that efficiency is then a big part of the puzzle on why the increased bandwidth to ram makes such an impact?
- sliken 5y agoMany things don't push the memory bandwidth and are cache friendly. However GPUs are bandwidth limited and the M1 Max does quite well against any other integrated graphics from Intel or AMD. Even on the CPU side it can be a big win, in particular on SpecFPRate (a collection of heavy floating point real world codes, not microbenchmarks) Anand has this to say: The fp2017 suite has more workloads that are more memory-bound, and it’s here where the M1 Max is absolutely absurd. The workloads that put the most memory pressure and stress the DRAM the most, such as 503.bwaves, 519.lbm, 549.fotonik3d and 554.roms, have all multiple factors of performance advantages compared to the best Intel and AMD have to offer. To drive this home compare the Spec2017 FP Rate, the M1 Max gets 81.07, the Ryzen 5950x (high end desktop with twice as many fast cores and a 105 watt TDP) gets 62.27. So a low power M1 Max with half as many cores and much lower power is 30% faster than AMD's highest end desktop chip. Instead of a desktop size/volume/power, you can get it in a laptop that's 2/3rd of an inch thick.
- klelatti 5y agoIt’s also very wide and has a big reorder buffer so ‘just slapping 16 gigs’ of cache probably isn’t the answer.
- errantspark 5y agoVery wide in the sense of memory bandwidth particularly, or is there another kind of wideness at play here? "has a big reorder buffer", I'm interpreting this as "one of the notable advantages of the M1 is it's ability to be more clever about processing instructions out of order to maximize resource utilization". Is that about right?
- klelatti 5y agoWidth in the sense of how many instructions can be executed at the same time. If you’re really interested in this it might be worth finding a copy of the Patterson and Hennessy book. It’s a big read and expensive (but older versions are on the internet archive [1]) and covers all these design issues in quite a lot of detail. [1] https://archive.org/details/ComputerArchitectureAQuantitativeApproach4thEdition https://archive.org/details/ComputerArchitectureAQuantitativ...
- errantspark 5y agoGotcha, I wasn't really aware of instruction level parallelism previously. I think the pieces of the puzzle are a lot clearer to me now, thanks for the replies.
- criddell 5y agoDo the gains you talk about include power consumption considerations? Could an Intel chip with a tone of cache run all day on typical laptop battery?