3 ms·
ARM has all these variations which make it seem as complicated as x86, but they are distict variations and future CPUs can for example drop 16 bit Thumb fairly
by TwoBit 5y ago
ARM has all these variations which make it seem as complicated as x86, but they are distict variations and future CPUs can for example drop 16 bit Thumb fairly clearly.
- api 5y agoIt's way easier to determine instruction length on ARM. It's usually fixed. That eliminates a lot of brute force thrashing that X86 decoders have to do. It doesn't impact transistor count all that much on a huge modern CPU but it saves a decent amount of power. It's one of the things that factors into why ARM is so power efficient. ARM has also been willing to drop older optional legacy stuff like Java oriented instructions that almost nobody used and Thumb. X86 supports nearly all legacy opcodes, even things like MMX and other obsolete vector operations that modern programs never use.
- thechao 5y agoIt's not that variable length is expensive, it's that variable length the way Intel does it is expensive. For instance — not that this is a good idea — you could burn the top 2b to mark instructions as 2/4/6/8 bytes (or whatever) in length. Then you can have your variable-width-cake-and-eat-your-fast-decode-too.
- richardwhiuk 5y agoVariable length is always going to restrict the parralelism of your instruction decode (for a given chip area/power/etc cost).
- thechao 5y agoOnly because a single instruction would take the "space" of multiple instructions in the I$-fetch-to-decode. The idea with variable-length encodings is that, for example, an 8B encoding does more than twice the work of a 4B encoding, so you lose out on the instruction slot, but win on the work done. I mean ... that's the theory.
- akira2501 5y ago> you could burn the top 2b to mark Which seems reasonable.. but you just burned 3/4 of the single opcode instruction space, which may not be worth it for most general purpose loads.
- bogomipz 5y agoWould you mind elaborating on the math of how the "top 2b" ends up burning 3/4 of the single opcode instruction space?
- brigade 5y agoEach bit used halves the potential encoding space, e.g. 2^32 -> 2^31 possible instruction encodings. For thumb, 32-bit instructions can be allocated up to 3/32 of the potential 32-bit space, and 16-bit instructions can use 29/32 of the 16-bit space (3 of the potential 5-bit opcodes denote a 32-bit instruction.) Which is probably a better ratio than 1/2 or 1/4 of each, for instance. Though I'm not sure how much of that encoding space is actually allocated or still reserved. Related, I believe ARM has allocated about half of the 32-bit encoding space for current A64 instructions.
- akira2501 5y agoFurther, if you want single-byte opcodes, then you took that space from 256 opcodes down to 64. It cost you 192 single-byte opcodes to use a 2b marker. This wouldn't be possible with the current x86 encodings [1]. [1]: https://www.sandpile.org/x86/opc_1.htm https://www.sandpile.org/x86/opc_1.htm
- bogomipz 5y agoAh OK, I think I understand now. You are specifically referring to the ARM Thumb instruction set as an example of this encoding scheme in both your comments?
- bogomipz 5y agoCould you elaborate - what is about the Intel design that makes the decode so inefficient? Is "2b" bits here? Are there examples of ISA or chips that handle variable length instruction encoding efficiently?
- Symmetry 5y agoIntel x86 isn't self synchonizing. In theory you have to decode every byte of the instruction stream that came before to be sure where the boundaries of an instruction are. Normally it stabilizes after a while in real world instruction streams but you can craft malicious instruction streams which yield two different valid sets of instructions depending on whether you start reading them at an offset or not. Contrast that to something like utf8 where that isn't possible.
- saagarjha 5y agoNote that random x86 code is usually “eventually self-synchronizing”, which id useful if you’re trying to disassemble a blob at a random offset.
- pkaye 5y agoDue to how the instruction set evolved, for the Intel x86 architecture, you have to look at a lot of bits of the instruction stream to determine where the next instruction starts. To execute multiple instructions per clock cycles you also have to decode instruction multiple instructions per clock cycle. I think this old Intel patent talks about one of their decoder implementations: https://patents.google.com/patent/US5758116A/en https://patents.google.com/patent/US5758116A/en
- atq2119 5y ago> you could burn the top 2b to mark instructions as 2/4/6/8 bytes (or whatever) in length. FWIW, this is exactly what RISC-V does.
- TwoBit 5y agoHow is this so? I thought RISC-V was fixed length, except for 16 bit compressed instructions. And afaik those aren't identified by a singular particular bit.
- atq2119 5y agoSee section 1.5 ("Base Instruction-Length Encoding") of the RISC-V spec. It's actually a bit more complex than just using 2 bits (I had forgotten those details), but the basic idea is the same in that there is a fixed cascade of bits identifying the instruction length. There aren't any standard extensions with instructions >32b yet, but the extensibility is there in the base spec.
- klelatti 5y agoI think that one of the things that distinguishes Arm from Intel (for good or bad) is that Arm _has_ left behind a lot of the legacy. aarch64 has no Thumb, Jazelle etc