3 ms·
> So another win for being able to read arm assembly. Yes, though that weird stuff with dollars in it is not normal AArch64 assembly! The article could have m
by bloak 1y ago
> So another win for being able to read arm assembly.
Yes, though that weird stuff with dollars in it is not normal AArch64 assembly!
The article could have mentioned the "stack moves once" rule.
- pjmlp 1y agoIt is due to the Plan 9 Assembly dialect most likely, because it wasn't enough that we already have differences between AT&T and Intel. https://go.dev/doc/asm https://go.dev/doc/asm Still, I find great that Go got back the 1990's tradition that compiled languages have an assembler as part of their tooling, regardless of the syntax.
- Neywiny 1y agoI've never heard of that rule (though tbh I'm not allocating > 64KB of stack when I'm in assembly) and it seems Google hasn't either. While I'm sure it makes sense, I don't think I've ever seen that be enforced. At least in C/C++. Maybe it makes more sense for these stack inspecting garbage collectors but I've also heard of ones that just scan the stack without unwinding anything. I did a test asking Google's AI to generate a complicated C function, put it in godbolt, and there's plenty of push push push push ..... Pop Pop Pop Pop going on
- rcxdude 1y agoDid you compile with optimisations? I think GCC will do a bunch of activity on the stack with -O0, but it'll generally coalesce everything into one push/pop per function with optimisations (not because of any rule, but just because it's faster). alloca and other dynamic stack allocation may break this, but normal variables should in pretty much all just get turned into one block on the stack (with appropriate re-use of space if variable lifetimes don't overlap)
- Neywiny 1y agoYes
- ori_b 1y agoIt will generate code to touch each page of the stack, because otherwise a very large stack allocation controlled by users (eg, in the case of a variable sized array) can be turned into a pointer to any location in memory by an attacker. Faulting in each page of the stack turns that into a crash. There was a userspace thread library I came across a long time ago that used variable length arrays to switch between thread stacks; the scheduler would allocate an array of the right size to bump the stack pointer to the different thread's stack.
- JdeBP 1y agoYou need to look at non-x86 architectures. It was common years ago on MIPS. * https://jdebp.uk/FGA/function-perilogues.html#StandardMIPS https://jdebp.uk/FGA/function-perilogues.html#StandardMIPS I wrote up the x86 equivalent of doing just two read-modify-write operations on the stack pointer over 16 years ago. * https://jdebp.uk/FGA/function-perilogues.html#Standardx86 https://jdebp.uk/FGA/function-perilogues.html#Standardx86
- mananaysiempre 1y ago> While I'm sure [bumping the stack pointer atomically] makes sense, I don't think I've ever seen that be enforced. At least in C/C++. That’s because the C ABI supports unwinding with a fairly expressive set of tools for describing stack-pointer state on a per-instruction level. Even the simpler Microsoft ABI essentially uses bytecode for that[1]; and on the more complicated Itanium ABI, you get DWARF CFI instructions, which make the correct way to preserve a(n x86) register in the function prologue look like push rbx .cfi_adjust_cfa_offset 8 .cfi_rel_offset rbx, 8 which are impossible to miss when reading compiler-generated assembly because of the sheer amount of annoying noise they create. The Go authors decided to sidestep all of this complexity, which is understandable to a degree, but apparently they did not think through all the ramifications of doing so. [1] https://learn.microsoft.com/en-us/cpp/build/exception-handling-x64 https://learn.microsoft.com/en-us/cpp/build/exception-handli...
- dwattttt 1y agoMS's ARM64 unwinding ABI looks even more complicated: https://learn.microsoft.com/en-us/cpp/build/arm64-exception-handling?view=msvc-170 https://learn.microsoft.com/en-us/cpp/build/arm64-exception-...
- mananaysiempre 1y agoEhh I wouldn’t say so (thanks for the correct link for ARM64 though in any case). What you need to be comparing to here is DWARF[1,2] section 6.4, and while it’s not as bad as other parts of DWARF, I still think it’s plenty complicated. [1] https://dwarfstd.org/doc/DWARF5.pdf#page=171 https://dwarfstd.org/doc/DWARF5.pdf#page=171 [2] Slightly modified by psABI[3] section 3.7 for x86-64 or the LSB[4] section 11.6 for ARM64, but at this point that’s a drop in the bucket as far as overall complexity is concerned. [3] https://gitlab.com/x86-psABIs/x86-64-ABI/-/jobs/artifacts/master/raw/x86-64-ABI/abi.pdf?job=build https://gitlab.com/x86-psABIs/x86-64-ABI/-/jobs/artifacts/ma... [4] https://refspecs.linuxfoundation.org/LSB_4.0.0/LSB-Core-generic/LSB-Core-generic.html#EHFRAMECHPT https://refspecs.linuxfoundation.org/LSB_4.0.0/LSB-Core-gene...
- 1y ago
- freep1zza 1y ago> Yes, though that weird stuff with dollars in it is not normal AArch64 assembly! See the AT&T vs Intel syntax since you aren't familiar with assembly: https://en.wikipedia.org/wiki/X86_assembly_language#Syntax https://en.wikipedia.org/wiki/X86_assembly_language#Syntax
- dpassens 1y agoThat's an x86 thing, though.
- indrora 1y agoThere are more assembler dialects than I care to remember. The 2A06 assembler that people who write NES code (and later on SNES/GB/etc) use has some real quirks: $ prefixes a literal hex value but % is binary, but # in front of that is an address, registers are baked into the opcode (ldx -> load into X), and more. Playstation folks all just used MIPS dialects which are mostly AT&Tish but the PS2 used an Intel style assembler.