4 ms·
Code density is an issue, requiring huge caches for Itanium to remain competitive. These large caches consume lots of die area, making the chips both expensive
by KMag 3y ago
Code density is an issue, requiring huge caches for Itanium to remain competitive. These large caches consume lots of die area, making the chips both expensive and power-hungry.
As I and others have pointed out, instead of a constant 3 instructions in 128 bits, a variable number of instructions in 64 bits should get you competitive code densities. (Say, have a 4-bit prefix that indicates if 1x60 bits, 2x30 bits, 4x15 bits, or 5x12 bits instructions are packed in the remaining 60 bits. For 12-bit instructions, you probably have the first instruction only able to store into r1 or r2, the second instruction only able to store into r3 or r4, etc. so that you have a 6-bit opcode, 1 bit for destination register and 5 bits for operand register. Plenty of old processors had an accumulator that was the implicit destination for most instructions.) Some of the other 4-bit prefixes would be used to indicate a combination of instruction widths and instruction parallelism to allow lower-powered in-order superscalar implementations.
Fixed alignment of 64-bit bundles of variable-width instructions makes parallel instruction decoding cheaper (and makes static analysis easier/finding ROP exploit targets very slightly harder).
Your high-performance cores are probably still going to have dynamic out-of-order scheduling in hardware. However, your power-efficient in-order cores might use parallelism data embedded in that 4-bit prefix.
Ideally, the hardware would have reservoir sampling for which branches are mispredicted and which instructions stall the pipeline. This information could be aggregated for a background process to re-optimize the binary, similar to what the current Android Runtime does.
Efficient hardware branch tracing (perhaps a second stack where, when the tracing machine state bit is set, the destination of each conditional/indirect branch target is pushed, along with an interrupt when the trace gets full) might allow for efficient runtime re-optimization of binaries, including inlining of dynamic library code into the code hot spots.
On a side note, I know RSIC-V at least at one point had a proposal for an extension to use the integer registers for floating-point. Does anyone have a feel for how costly it would be for register renaming logic to handle separate integer and floating point register files so that low-power implementations could use a single unified register file (perhaps without register renaming) and higher performance implementations (which would presumably have register renaming anyway) could use separate integer and floating-point register files?
In general, I hope we can find ISA designs that leave room for both very-low power implementations that can push some of the work into software, and high-performance implementations that aren't hindered by the features that allow lower-power/simpler implementations.