7 ms·
The complexity of current hardware is madness. The performance of software is influenced by the workings of tens of units (+local caches) connected (non-linear
by hydroreadsstuff 7y ago
The complexity of current hardware is madness.
The performance of software is influenced by the workings of tens of units (+local caches) connected (non-linearly) with fifos, ooo buffers, replay mechanism.
There are multiple versions of the same unit to save chip space (like light and heavy Integer/FP).
Units/domains have different clock speeds.
And from generation to generation port connections and instructions between pipelines can be reshuffled.
I suppose what keeps performance changes relatively straightforward for developers is the set of benchmarks used to evaluate the hardware early on.
I appreciate the article focusing on the I-Cache, and the nice intro to decoding.
I would have preferred having an example, and improving something instead of abstractly talking about problems, effects of code and optimizations and possible workarounds.
Tangent: I wonder if we will be seeing specialized instructions sometime that span multiple units and multiple cores in order to reduce data-movement. Think matrix-matrix multiplication. The potential improvements for power and speed seem huge.
- tempguy9999 7y ago> There are multiple versions of the same unit to save chip space (like light and heavy Integer/FP). I've literally never heard of this. I know you can have multiple execution units for parallel instruction execution, but 'light' and 'heavy' - can you give some info or a link? TIA
- hydroreadsstuff 7y agohttps://images.anandtech.com/doci/13699/Ronak26.jpg https://images.anandtech.com/doci/13699/Ronak26.jpg https://www.anandtech.com/show/13699/intel-architecture-day-2018-core-future-hybrid-x86/2 https://www.anandtech.com/show/13699/intel-architecture-day-... I believe this shows differences between FP and Integer units. In order to achieve a certain performance goal you don't necessarily need another integer divider when you want a new adder/multiplier. So you add a slimmer unit instead. I listed this in my original comment, because this is a giant can of worms for the compiler and decision-maker on where to execute what.
- tempguy9999 7y agoAh, thanks. The slide is interesting for extra reasons. With respect, I think you're misunderstanding. I thought you meant light/heavy versions of eg. adders, for some definition of light and heavy addition. I'm not an expert but... CPUs will put in extra execution units according to need (will typical code get faster with an extra X?) and cost. Shifters are typically very often used, and are simple. So are adders, though more complex. IIRC recent intel x64 will have several of of each[0]. Multipliers are less cheap so they have fewer (and often you can turn them into adds in certain cases such as progressive array lookups). Division is slow and very expensive in transistors, so they have 1 (division can often be turned into reciprocal multiplication anyway). Sqrt is even worse. And to repeat, I'm no expert and any corrections welcome. [0] <https://en.wikichip.org/wiki/intel/microarchitectures/coffee_lake> https://en.wikichip.org/wiki/intel/microarchitectures/coffee... If I'm reading this right, 2 shifters (2? I suppose they are fast so they are available soon after), 4 adders, 1 mult and 1 divider.
- deleted 7y ago[deleted]
- namibj 7y agoActually, barrel shifters are not that small. Added are significantly cheaper in terms of chip area.
- tempguy9999 7y agoSeriously?? It's the same bit pattern, err, shifted. Adders have got to carry, at each stage (ripple carry?). I am amazed, thanks.
- Robin_Message 7y agoThink of it this way: a barrel shifter has to be able to "carry" every bit to (potentially) every other bit.
- jandrewrogers 7y agoThis is, in a nutshell, why high-performance systems engineering is a rare skill set. It entails writing C++ (or whatever) with full understanding of the machine code that is likely to generate and how that machine code will interact with the incredibly complex internals of modern microarchitectures. It essentially requires de-abstracting two levels of abstraction below the programming language, which exist to reduce cognitive load, in your software design and implementation. It is unfortunate that this is still so useful in practice, given what it implies about the magnitude of waste in typical software systems.
- gumby 7y agoOne of the criticisms of the C++ standardization efforts is how many of the improvements are quitearcane. But those arcane features aren't necessarily for user code but to make library code efficient. Implementations of the standard library are often almost unreadable because of all the weird corner cases and performance optimizations they have to transparently handle. This can go a long way to improving user code...but of course it's no magic bullet.
- cbetti 7y agoAre you talking about STL interfaces or implementations?
- gumby 7y agoSTL in particular, but it's hardly the only case.
- adrianN 7y agoI really dislike this mindset in the C++ community. The people writing libraries are normal programmers too. They're not some kind of supermen on whom you can load infinite complexity. And it's not like the STL is this amazing piece of software that mere mortals can't beat. Quite the contrary, any sufficiently large piece of software tends to replace the STL at least in parts with their own implementation.
- namibj 7y agoYou mean Nvidia's tensor cores? You can buy what you wonder about for over a year now.
- legulere 7y agoAt least for binary matrix multiplication there is precedent: https://github.com/riscv/riscv-bitmanip/wiki/bitmat https://github.com/riscv/riscv-bitmanip/wiki/bitmat