5 ms·
"On the other hand, that wastes cache memory and fetch bandwidth..." This. One of the reasons why modern x86 processors have such strong performance is their e
by voidlogic 8y ago
"On the other hand, that wastes cache memory and fetch bandwidth..."
This. One of the reasons why modern x86 processors have such strong performance is their external facing CISC interface effectively acts as memory compression while they are RISC on the inside. In many ways its a best of both worlds that was achieved through incremental evolution.
- 0815test 8y agoNope, you don't need anything like the complexity of x86 decode in order to reach good instruction density. RISC-V matches x86 with a simple "compressed" extension involving 16-bit forms for some of the most common instructions (but still quite straightforward to decode in hardware).
- voidlogic 8y agoI agree. Its almost always more performant to intentionally build a system that to have characteristics that another system incrementally evolved to have. This is why rewritten software (if actually delivered) often performs better than the original, its usually this and not new framework X or new language Y that resulted in the win.
- phkahler 8y agoAnd RISC-V encodes programs into fewer bytes than X86, so it wins in that regard. But there are still no implementation that have all the other features needed for top performance. There are many factors that affect performance.
- voidlogic 8y ago"There are many factors that affect performance." Not only man different features contribute to performance but also that put performance depends so much on use-case. Two CPUs implementing the same arch might each win a benchmark that has a different instruction mix or memory access pattern, etc.
- AnIdiotOnTheNet 8y agoNot being any kind of expert in processor microarchitecture, I wonder if there is potential in just actually compressing instructions, as in with a huffman table. Could hardware decoding be made fast enough that the cache and fetch savings would make up for it?
- BeeOnRope 8y agoApple is already doing this (transparent memory/cache compression) on the GPU side, and there is some speculation they are doing it for the LLC on the CPU side, or may start doing it soon, based on patents they have filed.
- monocasa 8y agoModern ISAs like RISCV come up with their space conscious portions of their ISAs by taking a sort of human encoding perspective on the instruction stream. The definitely iterated on that base concept of allocate the number of bits needed based on frequency in the stream.
- pcwalton 8y agoNo, x86-64 is just as inefficient as AArch64 these days, because of all the REX prefixes. Almost every instruction needs one and it bloats the size tremendously, to the extent that most x86-64 instructions are just about 4 bytes. Measure binary sizes if you don't believe me.
- monocasa 8y agoI would love to see a newer CISC that didn't have all these required prefixes, and took a more Huffman encoding perspective so that instructions like hlt aren't allocated a single byte. Memory to memory ops essientially let you encode physical registers without using architectural registers, saving additional bits too.
- phire 8y agoI have encountered an instruction set which supported 16bit, 32bit, 48bit and 80bit encodings. Was basically Thumb2 but with more flexibility, and it still felt like RISC.
- userbinator 8y agoAlmost every instruction needs one and it bloats the size tremendously, to the extent that most x86-64 instructions are just about 4 bytes. That's only if your code happens to be particularly "64-bit-heavy", or the compiler isn't doing a good job at selecting registers; the original designers (at AMD, not Intel) decided on the prefices (and defaulting to 32-bit for most ops) instead of defaulting all operations to 64-bit in 64-bit mode precisely because it would be better for size and performance --- using their carefully optimised compilers. Plus, what can be done with a single 4-byte instruction on x86 can require multiple 4-byte ARM instructions, and that adds up quickly. I can't find it at the moment but one of the studies I remember comparing the binary sizes was using GCC, which is widely available and free, but probably one of the worst compilers at x86 size optimisation I've seen. I even recall a remark in that study about how it was generating mostly RISC-like instructions, so in other words they were comparing binaries generated for a RISC CPU using a RISC-oriented compiler with ones generated for a CISC CPU using a RISC-oriented compiler, failing to exploit the full capabilities of a CISC. I've written x86 Asm for several decades (started with 16-bit --- dating myself here...), and done some occasional MIPS and ARM, and it's very difficult for me to believe that the RISCs have any intrinsic advantage in code density other than the fact that compilers for x86 aren't that great at it; you can write a Fibonacci calculator for the latter in 5 bytes and pushes and pops are single-byte instructions, while on the former even a register-register move is 4 bytes.
- pantalaimon 8y agoIf you look at the benchmarks for x32 (ILP32) you see an improvement of up to 10% over x86-64. Granted this most benefits pointer heavy code, but is still doesn't make up for the lions share of performance.
- firethief 8y agoThat kind of benchmark shows the data path cost of big pointers but misses much of the instruction path cost of inefficient encoding, because the cost doesn't usually manifest as bottlenecking at the frontend. The cost is having a frontend that can keep up. On Intel cores this involves having a huge decoder that can deal with arbitrary alignment and unbounded sequences of prefixes; and having caches both for instructions and decoded uops (multiple types of the latter). Plus fundamental limitations on what the backend can do: there would be no point adding a 4th vector op x-port, because the frontend is miles away from being able to keep up with that many instruction bytes. All of that, and still the programmer/compiler walks a knife's edge avoiding frontend bottlenecks, trying to keep code tight and aligned so the important parts fit in loop buffer or at least uop cache and don't have to squeeze through the decoders. Use it or don't, REX is paid for.