4 ms·
Wow, if that's true then that is such a massive advantage. L1 is a single cycle fetch cache if it is like x86. So, individual cores can do compute so much bette
by actuator 5y ago
Wow, if that's true then that is such a massive advantage. L1 is a single cycle fetch cache if it is like x86. So, individual cores can do compute so much better by fitting more data at once.
- wmf 5y agoL1 hasn't been a single cycle for a long time, like decades.
- actuator 5y agoHmm, thanks; that seems interesting. I guess I need to read up more. On searching in Google someone[1] was quoting Xeon L1 fetch as approximately 4 cycles. I don't know if this is an average across branch prediction hit/miss, I will look for source of these numbers and try to read what has changed. [1] https://stackoverflow.com/a/4087331 https://stackoverflow.com/a/4087331
- moonchild 5y agoYes, 4-5 cycles is to be expected. Also, larger caches generally trade off latency, so I would not be surprised if apple's chips were slower still.
- 95014_refugee 5y agoThe cycle count is largely irrelevant in a speculative, pipelined core, as long as it can speculate far enough ahead that the delay is masked by other activity. 4-5 cycles is nothing in the scheme of things.
- moonchild 5y agoEmphasis on 'largely'. If you have e.g. multiply indirect pointers, then you care. That said, I didn't mean to imply that moving to bigger, slower caches was the wrong tradeoff.
- hajile 5y agoThis is part of what makes Apple's design so incredible. Normally, increasing the cache size dramatically increases latency, but they paid the cost and their 128kb D-cache and 196kb I-cache still retain 2-3 cycles of latency.
- nicoburns 5y agoYes, my understanding that one of the main reasons that Apple CPUs are so power-efficient is (probably) that they have absolutely huge caches.