4 ms·
> Using `lea` […] is useful if both of the operands are still needed later on in other calculations (as it leaves them unchanged) As well as making it possible
by pansa2 10mo ago
> Using `lea` […] is useful if both of the operands are still needed later on in other calculations (as it leaves them unchanged)
As well as making it possible to preserve the values of both operands, it’s also occasionally useful to use `lea` instead of `add` because it preserves the CPU flags.
- andrepd 10mo agoFunny to see a comment on HN raising this exact point, when just ~2 hours ago I was writing inline asm that used `lea` precisely to preserve the carry flag before a jump table! :)
- MYEUHD 10mo agoI'm curious, what are you working on that requires writing inline assembly?
- veltas 10mo agoI'm not them but whenever I've used it it's been for arch specific features like adding a debug breakpoint, synchronization, using system registers, etc. Never for performance. If I wanted to hand optimise code I'd be more likely to use SIMD intrinsics, play with C until the compiler does the right thing, or write the entire function in a separate asm file for better highlighting and easier handing of state at ABI boundary rather than mid-function like the carry flags mentioned above.
- vlovich123 10mo agoGenerally inline assembly is much easier these days as a) the compiler can see into it and make optimizations b) you don’t have to worry about calling conventions
- Someone 10mo ago> the compiler can see into it and make optimizations Those writing assembler typically/often think/know they can do better than the compiler. That means that isn’t necessarily a good thing. (Similarly, veltas comment above about “play with C until the compiler does the right thing” is brittle. You don’t even need to change compiler flags to make it suddenly not do the right thing anymore (on the other hand, when compiling for a different version of the CPU architecture, the compiler can fix things, too)
- veltas 10mo ago> “play with C until the compiler does the right thing” is brittle It's brittle depending on your methods. If you understand a little about optimizers and give the compiler the hints it needs to do the right things, then that should work with any modern compiler, and is more portable (and easier) than hand-optimizing in assembly straight away.
- andrepd 10mo agoWell in my case I had to file an issue with the compiler (llvm) to fix the bad codegen. Credit to them, it was lightning fast and they merged a fix within days. gcc optimised it correctly though.
- kragen 10mo agoIt's rare that I see compiler-generated assembly without obvious drawbacks in it. You don't have to be an expert to spot them. But frequently the compiler also finds improvements I wouldn't have thought of. We're in the centaur-chess moment of compilers. Generally playing with the C until the compiler does the right thing is slightly brittle in terms of performance but not in terms of functionality. Different compiler flags or a different architecture may give you worse performance, but the code will still work.
- EdwardDiego 10mo agoCentaur-chess?
- vardump 10mo agoMight be an interpreter or an emulator. That’s where you often want to preserve registers or flags and have jump tables. This is one of the remaining cases where the current compilers optimize rather poorly: when you have a tight loop around a huge switch-statement, with each case-statement performing a very small operation on common data. In that case, a human writing assembler can often beat a compiler with a huge margin.
- pedrocr 10mo agoI'm curious if that's still the case generally after things like musttail attributes to help the compiler emit good assembly for well structured interpreter loops: https://blog.reverberate.org/2025/02/10/tail-call-updates.html https://blog.reverberate.org/2025/02/10/tail-call-updates.ht...
- gishh 10mo agoI worked on a C codebase once, integrating an i2c sensor. The vendor only had example code in asm. I had to learn to inline asm. It still happens in 2025
- andrepd 10mo agohttps://github.com/andrepd/posit-rust https://github.com/andrepd/posit-rust LLVM codegen has been almost always sufficient, but for a routine that essentially amounts to adding two fixed-size bigints (e.g. 1024-bit ints represented as `[u64; 16]`), codegen was very very bad. Writing a jump table by hand literally made the code 3× faster :)
- AI-NoGuardrails 10mo ago[flagged]