7 ms·
I'm amazed at the features they've crammed into x86 without breaking compatibility; Definitely impressive. That being said haven't they perverted the whole indu
by MawNicker 10y ago
I'm amazed at the features they've crammed into x86 without breaking compatibility; Definitely impressive. That being said haven't they perverted the whole industry with this tactic? It seems like we ought to have switched from CISC to RISC. From a purely technical perspective at least. Microcode just seems like a run-time mini compiler. If it were just part of the actual compiler wouldn't that save power and heat? Could it hurt performance in some circumstances? Does CISC have any technical merit at all? I would love for this to have been motivated by something other than their x86 patents.
- erichocean 10y ago> Does CISC have any technical merit at all? Object code, in some cases, can be smaller. That's why modern RISC ISAs like RISC-V include a "compressed" instruction set option, to combat the object code size disparity with CISC ISAs.
- Unklejoe 10y agoI've heard of that. I believe it's used on some automotive computers (Bosch GS19 transmission controller comes to mind, which I think uses the Freescale MPC500). I wonder how it works. If the opcodes can all be compressed, why not just make them smaller to begin with? I guess I always thought of opcodes as being enumerations of "commands".
- wolf550e 10y agoRISC instructions are all the same size in bits to achieve fast decode speeds. Thus you must choose between 32bit instructions or 16bit instructions. ARM supports both, including function calls from one to the other.
- plasticchris 10y agoThis must have been solved, thumb 2 (https://en.wikipedia.org/wiki/ARM_architecture#Thumb-2 https://en.wikipedia.org/wiki/ARM_architecture#Thumb-2) mixes 16 and 32 bit instructions.
- gvb 10y agoThe 16 bit "thumb" instructions are still all fixed size - 16 bits. They do this by limiting the available options in the instructions. For instance, the register selection field will be 3 bits so only 8 registers can be directly accessed by the thumb instructions. They also drop some instructions IIRC. The 16 bit instructions map 1:1 to the 32 bit instructions, so decoding is unchanged, The difference is a pre-decoding transmogrificaiton stage where some of the resulting fields in the transmogrified 32 bit instruction are fixed (e.g. 0) and some are filled in from the original 16 bit instruction fields. Switching between 32 bit and "thumb" is fairly painless but the instructions cannot be intermingled. When you do a call, you have the opportunity to switch 32/thumb. When you return, the processor switches back as appropriate. The resulting thumb code is more compact than full 32 bit instructions, but it isn't 2x more compact because you still need some 32 bit instructions (subroutines) and the 16 bit instructions are not as powerful as the full 32 bit instructions.
- baobrien 10y agoThat's why I'm a fan of the RISC-V compressed ISA extension. Compressed instructions are 16 bits long and a subset of the 32 bit instructions, but can be mixed freely with 32 bit. The extension also relaxes the alignment requirements, so that 32 bit instructions can be aligned to 16 bit boundaries.
- oshepherd 10y agoThumb-2 (about 13 years old at this point) extended Thumb to cover the entirety of instruction set. That is, every* ARM instruction has a corresponding Thumb instruction (The converse is not quite the case) There is very little reason to use ARM code in ARMv7 and up. Incidentally, ARMv8's 64-bit mode (AArch64) adds a whole new instruction set, called A64. It's fixed width 32-bit per instruction, and the only option for 64-bit code. * Excluding some really obscure, mostly long deprecated ones
- oshepherd 10y agoI noticed that I missed out mentioning that, yes, Thumb-2 made Thumb variable length (it already sorta-was, branches were always kinda 32-bit) Any halfword where bit[15:13]=='111' && bit[12:11]!='00' is the leading half of a 32-bit Thumb instruction
- deleted 10y ago[deleted]
- baobrien 10y agoIn the case of RISC-V, the compressed extension only covers a subset of the full length instructions, but can be freely intermixed with them. It turns out that RV64GC (risc-v 64 bit with compressed instructions) usually beats x86_64 in dynamic instruction bytes. See slide 31: https://riscv.org/wp-content/uploads/2016/07/Tue1130celio-fusion-finalV2.pdf https://riscv.org/wp-content/uploads/2016/07/Tue1130celio-fu...
- djcapelis 10y agoI am not entirely disagreeing, but there are some interesting things one can do with a CISC processor. Arguably code density can be higher on a CISC device. You'll note that ARM has implemented basically code compression in their ISA with thumb. As for microcode on x86, most of the actual instructions running on the processor are pulled from L1 already translated. This used to be the trace cache, but after P4 that was dropped and now comes back as a u-op cache, meaning mostly that the instructions are stored in decoded format and so don't get run through translation again. (In P4's trace cache, branches were smashed and whole trace sequences up to three branches long were cached and then speculatively retrieved and executed.) Interestingly, there's research that shows you can spend plenty of time and gates to optimize the instructions as they come out of translation and go into the trace or u-op cache without negatively affecting performance, which allows some optimization basically for free. No one would do this these days because it's extra power and heat, but it was seriously considered back in the gigahertz wars. In addition, it should be noted that the translation between x86_64 and u-op is not super larger or challenging and is much much less challenging than most things we'd think about as a compiler pass. Decoupling internal architecture from externally exposed ISA is not an unreasonable choice, but x86 is certainly no longer a highly desirable external interface these days either. In a world where single thread performance isn't a significant differentiator, there's fewer reasons to use x86.
- freehunter 10y agoI can't answer the question of technical merit, but things hardly ever win public appeal based on technical merit. The appeal of x86 is the mountains of backward compatible software, which is due to how well-known its programming methods are, which is due to how incredibly popular the CPU type is. If we keep switching to the best possible CPU architecture every time a new one popped up, nothing would ever achieve what x86 has done. English is far from the best language, but everyone on this site speaks it. Gasoline is far from the best source of fuel for a vehicle, but every corner has a gas station. x86 is far from the best architecture, but basically every OS at leasts supports it in some way.
- mozumder 10y ago> It seems like we ought to have switched from CISC to RISC. From a purely technical perspective at least Uh, no. CISC is far more efficient than RISC. You can do much more with fewer bytes, increasing your cache efficiency. Also, the fastest instructions are just as quickly decoded as RISC, if not faster. And any complex instructions we can leave for microcode. The fact that early microprocessors are CISC shows you how much they could do with just a few thousand transistors. RISC architectures never took off until they could pack millions of transistors in, which should hint at how inefficient they were. This inefficiency manifests itself in modern times through power consumption. DEC Alphas & Power CPUS were already power consumption beasts, which is why Apple had to switch from Power to x86 for Macs. Right now there is nothing intrinsically better about RISC architectures. Like VLIW, RISC is a nice theoretical computer-engineering experiment, nothing more.
- gvb 10y agoMy take on the CISC/RISC "fight" is that the RISC was based on the assumption that the CPU clock was the limiting factor, so making a simpler CPU (RISC) would allow the CPU to be clocked faster than CISC and thus it would run faster. This was the case briefly when RISC came out (e.g. Alpha), but x86 fought back with improved fab technology and also threw more power at the problem. Lots of power. What blindsided RISC was that the CPU clock speed is not as much a limitation as memory access speed[1]. The voracious appetite of RISC instruction fetching[2] through the memory subsystem slowed it down more than it could speed up the CPU clock. (Yeah, yeah, broad generalizations.) Oh man, my point in a great graphic! https://www.cs.virginia.edu/stream/ https://www.cs.virginia.edu/stream/ [1] If you run the numbers, todays DDRn SDRAM latency is maybe half the latency of the original IBM-AT even though the advertised when streaming bandwidth is several orders of magnitude faster. In other words, if you are not streaming the memory accesses, your memories are not much more than 2x faster than the IBM-AT. Caching enables streaming. Branches break the stream. That is why both caches and branch prediction logic is huge. My numbers are out of date, but when I ran the numbers for the DDR2(?) memory in my hardware, the worst case latency was 56 clock cycles (e.g. if you had to close a page, open a new page, CAS latency, etc.). [2] RISC takes 1.5-2x more instructions than CISC in my experience. YMMV.
- detaro 10y ago> If it were just part of the actual compiler wouldn't that save power and heat? That gives less flexibility to the processor maker for individual optimizations. A CPU can optimize its microcode generation to use its hardware as efficiently as possible. Made a hardware change that can be used for more performance? Update the microcode generator and ship the new CPU, all updated (complex) instructions build on that are automatically faster/more efficient/whatever you optimized. If you compile to microcode-equivalent, your compilers don't know about the new detail, so they won't try to use it. So you need to ship a new CPU, and then update all compilers to know about your new microcode, and get software recompiled, and worst case shipped in both a version for your new and all old CPUs, ... Giving the CPU abilities to change things about how it exactly executes code means the compiler has to know less about the specific CPU running the code later. AFAIK this is one of the things that killed the Itanium line: they tried to rely on sufficiently smart compilers instead of hardware optimizations, and had a hard time to adapt the code running to innovations on the hardware level.
- icebraining 10y agoThere's an alternative: running an optimization step after software installation, like Android's ART does.
- dbcurtis 10y ago> It seems like we ought to have switched from CISC to RISC. If you ever have a question about computer architecture, ask two questions: 1) where does the memory bandwidth go, and 2) where does the die area go. All else follows from this. RISC only made sense when CPU clock speeds were at rough parity with central memory speed, on-chip I-caches were limited, pipelines were short, and compilers were only good at modest optimizations. If you look at functionality per instruction byte, RISC is fairly low. In today's world, CPU clocks are hugely faster than memory speeds, on-chip I-Caches are huge, pipelines are very very deep, branch prediction hardware is very good, and compiler back-ends are much much smarter than the peak days of RISC. CISC wins because the amount of functionality that you can move into the I-Cache per clock is higher and the amount of functionality you can keep in a given amount of I-Cache die area is higher, and compilers are good enough to target specialized instructions efficiently. Is what Intel has crammed into X86 impressive? Yes, very. And it was damn hard work. I can tell you that walking all the way out to byte 15 of an instruction to look at the MOD R/M byte to decide if you have to raise an illegal instruction exception is a painful long path to squeeze under the clock constraint. But breaking backwards compatibility is just not something customers will put up with. So logic and circuit designers get to ply their trade in it's most convoluted form with the X86. Perhaps these days I should amend my question list and add a third: 3) where does the power go? This is where ARM has an advantage over X86. The equations that you have to resolve to issue an X86 instruction are simply more complex than for ARM, and all that bit-flipping consumes power.
- rayiner 10y agoRISC versus CISC was a non-issue even 10 years ago. X86 was never that CISC-y, and x86-64 is straight up a pretty clean architecture, with a memory model that doesn't suck.
- MawNicker 10y agoToo pithy for the down-voters apparently. This comment made me laugh. What x86 taught me is that: CISC = RISC + microcode.
- legulere 10y agoThe original x86 was pretty unriscy with the registers not being general purpose. The addressing modes to this day are still very CISCy
- fanf2 10y agoThe addressing modes in x86 are much less CISCy than the VAX or 68040 - no multiple indirections or increments in one instruction. John Mashey explained well the difference http://yarchive.net/comp/risc_definition.html http://yarchive.net/comp/risc_definition.html
- fanf2 10y agoThe addressing modes in x86 are much less CISCy than the VAX or 68040 - no multiple indirections or increments in one instruction. John Mashey explained well the difference http://yarchive.net/comp/risc_definition.html http://yarchive.net/comp/risc_definition.html
- pcwalton 10y ago> x86-64 is straight up a pretty clean architecture, with a memory model that doesn't suck I wouldn't go that far. In particular, the encoding of instructions in x86-64 makes zero sense except for backwards compatibility. Since all code is littered with REX prefixes, you have the same i-cache footprint as RISC architectures do without the benefits (die area for decoding, etc.). If you fixed the instruction encoding, extended the three-address forms of instructions introduced in AVX2 to scalar integer operations, and finally threw away real mode and port mapped I/O, etc., I agree that x86-64 would be pretty nice. I'd be happy to get even one of those improvements, honestly :)
- Symmetry 10y agoDoes CISC have any technical merit at all? RISC and CISC are both philophies that rose out of the contstraints of the eras they originated in and both, in their pure forms, are totally obsolete. The effort that goes into the design of a modern CPU micro architecture is so high that the number of instructions has a pretty small impact on the overall design effort. The descendants of the original RISC ISAs have kept adding new instructions and for good reason. On the other hand part of the increase in micro architectural complexity is that chips are designed to decode and issue multiple instructions per clock cycle. So the ease of multiple decode where every instruction is aligned to 32 bit boundaries is actually a noticable advantage. The size of the instruction stream is still an issue but on that front x86-64 and ARM-64 are more or less tied these days in terms of bytes of instruction per task so that's a wash in practice these days.