4 ms·
In other words, READ THE ASSEMBLY!
by Validark 2y ago
In other words, READ THE ASSEMBLY!
- dijit 2y agothe assembly can mislead you, i have seen complex, long assembly that even does memory reads perform 2.5x the performance of “simple” assembly that would have been 4-5 lines. Benchmarking is the only thing you can really do in some cases.
- neonsunset 2y agoThis is an edge case. Disassembly may not tell the whole story but it is the definitive answer to what the compiler thinks about your code. Given enough knowledge it is not difficult to account for cache/memory accesses, see where store-to-load forwarding may or may not kick in, consider optimal independent operations scheduling, check if the code is too branch heavy or not, etc etc. And once you're done, you can do another round with tools which can do more detailed analysis like https://uica.uops.info https://uica.uops.info I do this when writing performance-sensitive C# all the time (it's very easy to quickly get compiler disassembly) and it saves me a lot of time.
- Validark 2y ago> Benchmarking is the only thing you can really do in some cases. I think the complete opposite! Benchmarking is very difficult to get right and in some situations you couldn't really make a benchmark to test what you're interested in improving. Reading assembly can be objective if you know or look up the instruction latency/throughput of each instruction (or you have a tool provide it) and you can also use loop throughput analyzers (as the other commenter mentioned) that will try to predict the typical throughput for a given loop. IMO if you get used to looking at assembly it becomes obvious in the majority of cases whether there's performance left on the table. And "simplicity" does not equal "faster". Although for cold code, reducing its impact on i$ and d$ of the rest of the system is probably smart, so sometimes speed is not the only factor.
- menaerus 2y ago> the assembly can mislead you If only metric you look for is number of instructions then yes and this is a wrong way of looking at this type of problems. One can write a 50 LoC heavily optimized SIMD code that outperforms the compiler generated 10 LoC by 5x.