4 ms·
Well, it means that if code cannot fit everything into the ISA registers, it has to spill them to the stack, which AFAIK is not renamed to the physical register
by devit 11y ago
Well, it means that if code cannot fit everything into the ISA registers, it has to spill them to the stack, which AFAIK is not renamed to the physical registers on current x86, so you still need to pay an extra penalty to access the stack data in the L1 cache.
- Tuna-Fish 11y agoBut stack accesses are generally perfectly predicted, meaning that your loads from stack get executed and the data loaded into those extra registers long before your code needs to use those values.
- tux3 11y agoThe prefetcher does a great job, but it's not remotely enough to make stack access penalties disappear. It's trivial to write microbenchmarks where spilling hot read/write variables to the stack destroys performance, much harder to find special cases where the difference isn't noticeable.
- stephencanon 11y agoThe hot region of the stack is, for practical purposes, always resident in D$, so there's no prefetching to be done. I think that you're really talking about out-of-order and speculative loads. You're absolutely correct that spill/fill can be catastrophic for performance, however.
- stephencanon 11y agoThere are two common scenarios where spill/fill has catastrophic performance implications: - Workloads that are LSU-bound, like simple per-pixel image operations and level 1 and level 2 BLAS[1]. Here every spill is taking up two LSU dispatch slots that would otherwise be used for "real work". - Spilled loop-carried dependencies in tight loops; here you're simply adding latency to the critical path. Both are fairly rare in "generic" code, but very real concerns for compute kernels. [1] The BLAS operations have few enough buffers that spill/fill never actually happens in practice, but more complex operations on multiple buffers do run into this.