7 ms·
I never felt more like stabbing myself than when trying to cipher out exactly which immediate values are possible on ARM, and which are not. X86 I happen to en
by honkhonkpants 10y ago
I never felt more like stabbing myself than when trying to cipher out exactly which immediate values are possible on ARM, and which are not. X86 I happen to enjoy. It is not "the worst ISA" by any means. It has wonderful code density, which turns out the be very important. There's a reason that x86 won and continues to win.
- JoeAltmaier 10y agoWhen originally invented the x86 instruction set was efficient - the most-used instructions had shorter byte code sequences. But eventually some instructions got 'left behind' by the compilers. There are a whole host of single-byte instructions that are never, ever used by a compiler - the register exchange instructions for instance (xchg eax, ebx). Compilers just schedule destination registers carefully, never need to swap them around. Also the whole set of exchange-register-with-itself were defined but never used. E.g. xchg ax,ax which does nothing in one byte. In fact that one was considered useful, its used as the 'no-op' instruction (0x90) right? But what about xchg bx,bx, xchg cx,cx and so on? Just wasted single-byte opcodes. Leaving actual common instructions to use longer bytecode sequences. So maybe an executable should begin with an opcode-decode-table that is loaded with the code, that tells the hardware what byte sequences mean what instructions. So each executable code can be essentially compressed, using optimum coding for exactly the instructions that code uses most often. Just thinking out loud.
- qwertyuiop924 10y agoThat would be cool, but I would not want to hand-code assembly on such a platform. That makes memory segmentation look fun.
- JoeAltmaier 10y agoThe table could be created from your assembly automatically? I would never want to code the hex directly in any case.
- qwertyuiop924 10y agoThat could work... I thought it couldn't a second ago. I don't know why. Makes no sense to me now.
- wolfgke 10y ago> E.g. xchg ax,ax which does nothing in one byte. Luckily in x86-16 xchg ax,ax simply is encoded as 0x90, which is the same as nop (the same holds for xchg eax, eax in x86-32). > But what about xchg bx,bx, xchg cx,cx and so on? This (cleverly?) cannot be encoded in one byte. Here you have to use at least two bytes (0x87 followed by the ModR/M byte in x86-16; the same holds for their 32 bit counterparts xchg ebx, ebx; xchg ecx, ecx etc. in x86-32): > http://x86.renejeschke.de/html/file_module_x86_id_328.html http://x86.renejeschke.de/html/file_module_x86_id_328.html --- ADDITION: > So maybe an executable should begin with an opcode-decode-table that is loaded with the code, that tells the hardware what byte sequences mean what instructions. The engineers of the Transmeta Crusoe/Efficion processor tried something similar: > https://en.wikipedia.org/w/index.php?title=Transmeta_Crusoe&oldid=742538653 https://en.wikipedia.org/w/index.php?title=Transmeta_Crusoe&... "Crusoe was notable for its method of achieving x86 compatibility. Instead of the instruction set architecture being implemented in hardware, or translated by specialized hardware, the Crusoe runs a software abstraction layer, or a virtual machine, known as the Code Morphing Software (CMS). The CMS translates machine code instructions received from programs into native instructions for the microprocessor. In this way, the Crusoe can emulate other instruction set architectures (ISAs). This is used to allow the microprocessors to emulate the Intel x86 instruction set. In theory, it is possible for the CMS to be modified to emulate other ISAs."
- spc476 10y agoBeen done (somewhat). The PERQ (https://en.wikipedia.org/wiki/PERQ https://en.wikipedia.org/wiki/PERQ) had a writable instruction set. And I think Transmeta was trying to do something similar with translating x86 code to an internal format for execution.
- honkhonkpants 10y agoAre you using the past tense because you are going to tell us about another contemporary ISA that has more code per byte? X86 _is_ compact and this continues to be relevant to performance today.
- userbinator 10y agoIt is still efficient for general-purpose code compared to other ISAs; the majority of register-register and register-memory ops are 2-3 bytes, while for something like MIPS no instruction is ever shorter than 4 bytes. Compilers just schedule destination registers carefully, never need to swap them around. Actually, it can happen with certain instructions that need fixed register constraints (multiply, divide, string ops) --- I've encountered a few cases where, had the compiler knew about the exchanges, it could've avoided using another register or spilling to memory. As far as I know, in modern x86 cores the reg-reg exchanges are handled in the register renamer, so they aren't slower than using an extra register and definitely faster than spilling to memory (which might happen anyway for something else if it needed the extra register.) To witness, here is something no compiler (software) I know of can generate, even when given code that could generate it: theloop: xchg eax, edx add eax, edx loop theloop Never say never ;-)
- qwertyuiop924 10y agoWell, yes. That part sucks. But on x86, everything is like that. I'd rather have one weird, inconsistant thing than an amorphous, ever-shifting mass of them. Which is x86 in a nutshell. Here's a handy heuristic: if somebody claims to know every x86-64 instruction (or even every x86-32 instruction), you can be at least 90% sure they're lying.
- lgeek 10y ago> Here's a handy heuristic: if somebody claims to know every x86-64 instruction (or even every x86-32 instruction), you can be at least 90% sure they're lying. I very much prefer ARM myself, but you can probably apply the same rule of thumb to it. There are upwards of 400 instructions both in AArch32 and AArch64, with a fair number of differences between the two. Edit: I've also posted a breakdown of the immediate limitations for ARM in this thread. It's not that complicated when sticking to the standard instructions.
- qwertyuiop924 10y ago...But how many of those instructions are just variants of other instructions, but with a different addressing mode?
- lgeek 10y agoFair enough, some instructions have many variants. But I work with ARM assembly almost daily and I still wouldn't remember that, for example, 'VQRDMULH' is a real instruction and it stands for 'Vector Saturating Rounding Doubling Multiply Returning High Half'.
- qwertyuiop924 10y agoAh, yes. A fun one. Those exist in x86, too. I should probably count up x86 against ARM, so I have more than guesses to go on here. Maybe that part of x86 actually better.
- lgeek 10y ago> I never felt more like stabbing myself than when trying to cipher out exactly which immediate values are possible on ARM, and which are not. It depends if you're compiling for ARM or Thumb. The main rules are: ARM and 32-bit Thumb instructions: 12 bits for arithmetic data processing instructions and load/store, 8bits with even rotate right for bitwise data processing instructions + MOV and MVN 16-bit Thumb: 3 bits for arithmetic instructions with Rd != Rn, 5 bits for shifts (so the whole range is covered) and load / store (but shifted left by the size of the data, i.e. the maximum offset for this version of LDR is (0x1F << 2) = 124), 8 bits for arithmetic instructions with Rd == Rn, SP-relative loads / stores and literal loads. I doubt x86's popularity has much to do with its instruction set at this point. In particular, the variable length instructions are a pain for decoders (both hardware and software).
- ris 10y ago> There's a reason that x86 won and continues to win. It is definitely, definitely not the ISA.
- honkhonkpants 10y agoOk. It is though.