5 ms·
Oh for pete's sake, the die size cost for supporting the legacy x86 instructions is a fraction of a percent on a modern multicore CPU. This is because the trans
by jonstokes 9y ago
Oh for pete's sake, the die size cost for supporting the legacy x86 instructions is a fraction of a percent on a modern multicore CPU. This is because the translation hardware stays relatively fixed (in terms of size and complexity and transistor count), while Moore's Law keeps adding more hardware to each core and more cache, etc..
Take a look at this annotated die shot of an old Pentium 4:
https://i.stack.imgur.com/QK4gm.jpg https://i.stack.imgur.com/QK4gm.jpg
Up at the top is the microcode memory and microcode sequencer. That's the hardware that's responsible for translating all those old, large legacy instructions. It's not really much space on the P4, and the P4 Northwood is a 55-million transistor CPU.
Nowadays, depending on if you're talking about a mobile part or a higher-end desktop or server part, the CPUs have between ~1.5 billion to 2 billion transistors.
Again, that microcode ROM just doesn't grow very much as you add instructions, even a ton of instructions. And this is, well, one reason why Intel just keeps adding instructions. It's practically free.
- STRML 9y agoThat's a great graphic! I see a few more from around 2003 on http://chip-architect.com/ http://chip-architect.com/, do you know of anyone who's doing modern chips (x86 or ARM) this way?
- api 9y agoYou're right when it comes to die size and transistor budget, but complexity can sometimes have hidden costs. Complexity can limit and constrain the design of other aspects of a system in ways that are nearly impossible to predict and often non-obvious even when you're down in the trenches. Not only can today's features limit tomorrow's design, but because those features give rise to mental models and habits of thinking they can limit future design in an almost unconscious way. The more complex your system becomes the more its existing complexity dominates your thinking to the exclusion of innovative ideas. This process happens so quickly in software that it's obvious there: http://www.retrologic.com/jargon/R/rococo.html http://www.retrologic.com/jargon/R/rococo.html
- Retric 9y agoIt's not just ancient instructions that add costs, P4's MMX is also outdated etc. Further the core issue is the cascade where A limited B and C is limited by B directly and A indirectly, now continue that though 15 generations and things get complicated.
- tambourine_man 9y agoHey John, your in depth articles on Ars are missed. Ever thought of getting back on writing them occasionally?
- CalChris 9y agoI wish Intel (IACA, VTune, ...) or Agner Fog would list instructions which divert to the MS-ROM. A 2 or 3 instruction equivalent is often faster.
- Symmetry 9y agoPeople frequently complain about the complexity of decoding the x86 ISA and that seems to be what you're arguing against. But I didn't see the OP as making that argument at all. Instead it's things like the way x86 handles integer condition codes that has a huge effect on the complexity of the reorder buffer because now you have to track this additional input for each of your instructions. But quite apart from the transistor cost of various legacy feature is the engineering cost. When designing a new x86 chip you have to make sure that every instruction works correctly in every mode that supports it and that's hard. NRE is a large fraction of the cost of each chip and verification is a big part of that. Intel had all the transistors in place to support SMT early on in the Pentium IV development but they didn't turn it on until Northwood because it was hard for them to make sure it worked reliably. And recently look at the problems Intel has been having getting TSX to work right or the problems AMD had with legacy modes in Ryzen. And also, decode might not take very many transistors but those transistors are in constant use and consume way more power than the huge seas of 8T SRAM cells in the local cache despite being hugely outnumbered. This is far from the majority of the chip's power but every little bit counts.
- dtech 9y agoIs 1 transistor always 1 transistor? Afaik most of the increased number of transistors and die size comes from larger caches. I could imagine the translation hardware having a much larger impact on cost, complexity and performance than yet another few million transistors of L3 cache. Note: I have zero experience in hardware design
- xorblurb 9y agoThe whole impact of the architecture on the processor is to be considered, not just the size of the decoder. If the ISA was irrelevant (and the ISA does not just influence the decoder, it influences pretty much everything in various degrees) in modern days and even semi-modern days, Intel, with all its power (vastly superior to Arm, from a financing pov), would have come with something that does not suck in mobiles given how strategic this market is. It has not. In other words, you should not confuse the size of the lexicon with its associated comprehensive semantics. It does not matter that adding instructions in that framework is not too costly. The framework itself is costly for some applications...
- wolfgke 9y ago> Intel, with all its power (vastly superior to Arm, from a financing pov), would have come with something that does not suck in mobiles given how strategic this market is. This market is a lot more price-sensitive than the desktop, laptop and server CPU market.