4 ms·
Unless you want programmers to explicitely manage this fast memory the hardware is a better place to do those optimizations than the software. Compilers (even
by obl 10y ago
Unless you want programmers to explicitely manage this fast memory the hardware is a better place to do those optimizations than the software.
Compilers (even the high-end modern ones) are very quickly clueless about both memory and execution profiles. Those problems are hard to reason about statically. Is this pointer the same as this other one ? Will this loop be executed 10 times or 10^6 times ?
The hardware being a dynamic optimizer has a lot of runtime information (pointer values, cache entry usage, ...) to take advantage of.
Sure, one can argue that we just need better JITs, PGO, and programming languages that are easier to analyze statically. In the current state of affair we are way better off leaving this to the hardware.
By the way you do have "some" control over it using NT moves on x86. Turns out almost nobody uses them except by hand in very specific peak-bandwidth code. Compilers don't dare emitting those since the penalty of getting it wrong outweights the benefits of getting it right.
- restalis 10y ago"In the current state of affair we are way better off leaving this to the hardware." That is something I'm not so sure about. That might have been true before, when the hardware served well the then simple computational model of a single high frequency processing core. Things got more complicated since the times of first CPU cache offerings and hardware general use-case implementation can only give you so much. The performance boost that once could be relied on without much care now dissipated. This cache thing now become just a leaky abstraction. We can not abstract it away completely if we care for performance, nor we can control it in any meaningful way. It is not a good model anymore. Just recognize it as a failed experiment. "Unless you want programmers to explicitely manage this fast memory..." Yes, that's exactly what I have in mind (and implied before). Most applications do not need maximal performance and can disregard the fast memory completely. That will only mean that those processes either will not consume their share and leave it more for the processes that need it, or that the fast memory will (occasionally) be used indirectly, through the optimized bits in the layers that the given process relies on. This memory, however, won't be unproductively overwritten or invalidated on every cold cache or whatnot.
- dman 10y agoI use non temporal stores all the time when writing database related code.
- gaius 10y agoUnless you want programmers to explicitely manage this fast memory That's how we did it in the good old days on the 6502 (zp address mode) and we liked it like that...