4 ms·
Having looked at a fair amount of AT&T syntax assembly code over the years, the sheer nonsensicalness of the AT&T syntax cannot be understated. You seriously w
by csense 6y ago
Having looked at a fair amount of AT&T syntax assembly code over the years, the sheer nonsensicalness of the AT&T syntax cannot be understated.
You seriously want to argue that disp(base,index,scale) is clearer than [disp+base+index*scale]?
Is there any excuse, other than sheer laziness on the part of the parser writer, to require EVERY NUMERIC CONSTANT to be prefixed with a # and EVERY REGISTER to be prefixed with the percent sign?
So I can confidently say that the purpose of AT&T syntax was to make assembly programmers want to gouge their eyes out.
- nn3 6y agoFor the registers the % prefix avoids needing to add some prefix to every symbol. Otherwise you couldn't have symbols with the same names as registers. Not sure what you mean with every numeric constant prefixed with #. That's not the case on x86 AT&T assembly. Immediates require $, but that's the same as the Intel syntax.
- jwilk 6y ago> Immediates require $, but that's the same as the Intel syntax. No, "$" is not needed (or even allowed) for immediates in Intel syntax.
- adrian_b 6y agoAny assembly language needs to distinguish between immediate values and addresses. So you need a prefix on one of them, unless you use different mnemonics for the same operation, depending on whether it uses an immediate value or a memory operand. Doubling the number of operation mnemonics is clearly worse than having a single prefix. The prefix should be required for the less frequent case, and that is normally the use of immediate values. The Intel syntax as seen in the Intel architecture manuals is somewhat simplified, because it does not contain symbols and expressions. The complete assembly language syntax, for example the one implemented by the Microsoft MASM, needs frequently to distinguish the immediate values. Because MASM followed the Intel manuals, it did not use a prefix on immediate values but instead it used ugly and verbose prefixes on the memory operands, e.g. "WORD PTR ". In order to avoid excessive verbosity, MASM omitted the memory operand prefix in many cases where immediate values could not be used, which is possible only because the Intel instruction set is much less orthogonal than almost any other instruction set. Always using a prefix to distinguish immediate values from addresses is preferable to the inconsistent MASM method, because it can be tricky to remember which instructions can or cannot have both immediate operands and memory operands, so that the prefix is mandatory. Moreover, later ISA extensions could also change the available addressing modes for an operation, possibly mandating the use of the prefix. In conclusion, the MASM way of distinguishing the immediate values, by using "WORD PTR ", "BYTE PTR " and other similar prefixes on many but not all of the memory operands is certainly worse than always using the prefix "$" on immediate values.
- gliptic 6y agoWORD PTR etc. are size prefixes that would be needed even with some $-prefix for immediate values. I don't see the connection. Immediates and memory operands are already distinguished by []. The equivalent to WORD PTR etc. in AT&T syntax is the size-suffixes on instructions (addl, addw etc.).
- exmadscientist 6y ago> Otherwise you couldn't have symbols with the same names as registers. Some of us would argue that's a feature, and a very important one at that!
- tlb 6y agoThere are a lot of registers. 150+ named ones, including rarely used ones like SPL -- the low 8 bits of the stack pointer. This restriction would break existing assembly code whenever Intel added a new one, which happens regularly.
- inkyoto 6y agoAdvantages of having a common set of conventions might not be apparent if you look at them through the lens of a single ISA (x86 in your example); however, if you consider that porting UNIX and a C compiler to a new hardware architecture was everyone's favourite pastime back in the day, having the $ and % conventions greatly helped with reducing the cognitive load and shielding the engineer from quirks and idiosyncrasies of new assembly languages / ISA's. Registers can have vastly different names across hardware architectures, e.g. they could be named r0, g0, gr1, f2, Motorola 68HC16 has registers named X and Y, or simply numbered register names have been used as well (if my memory serves me well, the IBM System 360 ISA was the most egregious example where LD 1,1 meant load a constant of «1» into the register number 1, albeit I can't remember in which the order). Then there are widly varying processor specific control register names, e.g. i860 had «psr», «epsr», «db», «dirbase», «fir», «fsr», «kr», ki», «t» and «merge» control registers in the ISA, whereas the PDP-11's own processor status word (PSW) was mapped to the 177776 address in the hardware and PSW was a widely used (but not mandated) MACRO-11 preprocessor alias that had to be manually defined in every assembly file. Also, the register window SPARC architecture utilises names gN for global registers and lN for local registers, and with C compilers habitually emitting jump label names that start with a «L», without the help of extra semantic cues, it would not be immediately obvious whether ld l0, l0 meant «load the address designated the label named l0 into a register named l0», or «store the contents of the register named l0 into an address the label l0 refers to» (I am stretching the SPARC example a bit here to highlight my point as it has a load/store architecture). Therefore seeing a $ was an immediate «heads-up, a constant is about to follow» and, likewise, seeing a % indicated a upcoming register name. It certainly has helped me quite a bit with squickly settling into a new architecture by skimming over assembly outputs of C compilers and assembly parts of UNIX kernels. AT&T conventions were rather widely applied across for some time, but it appears that they are less frequently followed nowadays.
- stormbrew 6y agoI'm honestly ok with prefixing registers with a consistent symbol, and discerning for eg. immediates and literal addresses in a uniform way is pretty useful in a lot of instruction sets. I think those are ok. On the other hand, though, there are the ways that it diverges semantically from the platform's norm: Indexed-mode, like you said, adding suffixes (for word size) to opcodes, the always popular reversal of src,dst arguments. That's where it bothers me. When you put it all together though it just gets a lot more frustrating. On the plus side at least you always know what you're looking at.
- Someone 6y ago“You seriously want to argue that disp(base,index,scale) is clearer than [disp+base+indexscale]?”* Did anybody ever argue that?I always thought the AT&T code was indicative of a meta-assembler (https://www.encyclopedia.com/computing/dictionaries-thesauruses-pictures-and-press-releases/meta-assembler https://www.encyclopedia.com/computing/dictionaries-thesauru...: “meta-assembler A program that accepts the syntactic and semantic description of an assembly language, and generates an assembler for that language. Compare compiler-compiler.”) Such a meta-assembler would not know that this instruction does a few adds and a multiplication. It just would know how to pack a few values of varying bit width into a few bytes to output. And yes, a meta-assembler could add syntactic sugar to support that, but they date from a time where kilobytes were expensive, and such ‘fluff’ deemed not worth it. That also may be why it uses prefixes for register names. It avoids lots of relatively costly lookups in a symbol table. It also may make it easier/feasible to write a one-pass assembler.