3 ms·
I want to really love this post. But I am not smart enough to know why it’s awesome. Anyone care to give a digest for mere mortals?
by mooreed 7y ago
I want to really love this post. But I am not smart enough to know why it’s awesome. Anyone care to give a digest for mere mortals?
- fundamental 7y agoThis tool helps look at very small chunks of assembly code to identify the cost via a simulated execution. Beyond that it focuses on multi-stage execution of instructions in a hardware pipeline, providing tools to show the execution as well as flag which factors would be the limiting factor for the code in question. The level of detail that it provides as well as the apparent simplicity of using the tool makes it unique (from what I've seen). Understanding where the limiting factor is can help optimize and tune small chunks of code.
- tux3 7y agoOh it really is an awesome low-level tool, but it's not as complicated as it sounds! Here's my attempt at explaining it, hope this helps :) The gist of it is that your CPU loves to multitask (because that's so much faster), and if you like performance you want to maximize how many things it's doing at once at any time. You want every part of the CPU to have work on its schedule all the time so it doesn't sit idle. This shows you what the schedule for your code looks like so you can optimize it. -- In more details, this tool is going to read your code instruction by instruction, compute what kind of schedule the hardware will be able to make for executing each instruction, and tell you which circuits might be overworked or going unused. (Caveat: the CPU makes its own schedule on the fly as best as it can — this tool is just an approximation in software — but your compiler's best guess is still a pretty good guess in general!) For example (actual numbers! [0]) your CPU could have 4 "ports" (circuits) that can start doing math on integers (with their ALUs) at any instant in time, or 2 of those "ports" could also multiply floating-points numbers, but each circuit can only be given one kind of task at a time to keep things manageable. Well it turns out making a good schedule is a surprisingly hard problem, since if you send int work to the first two ports and you didn't foresee there'd be float work coming after, the first two ports will have a full schedule while the other two will be doing nothing. A better schedule could have had all four of them busy in parallel! You don't directly have a say in the schedule — the hardware does its best — but with LLVM-Mca (or Intel's IACA, which inspired it) you can write code that you know will be easy to schedule in parallel, and that's already some pretty awesome tools to have! [0]: https://en.wikichip.org/wiki/intel/microarchitectures/skylake_(client)#Scheduler https://en.wikichip.org/wiki/intel/microarchitectures/skylak...
- jcranmer 7y ago> You don't directly have a say in the schedule — the hardware does its best That's not quite true. The order you present the instructions influences the actual execution order greatly, and it's why instruction scheduling remains an important part of compilers.
- tux3 7y agoYep, I wasn't sure how to word it (it's pretty hard to keep it short without being too wrong!). You're right that there are many many things you can do to the code to influence the scheduling — like the reordering the compiler is doing — and at the end of the day that has a predictable impact on the scheduling. Don't get me wrong I love my compiler, and the fact that we can impact scheduling is why LLVM-Mca is useful in the first place. What I meant to write is that your x86 isn't some kind of mostly statically scheduled VLIW. the behavior of the hardware is only partly predictable, and even IACA has to make some tragic simplifications. Tweaking the alignment to play with fetch boundaries has an effect, vectorizing obviously does, picking a different mix of instructions can help, artificially loading a port to prevent a bad scheduling decision down the line is not always entirely stupid, etc... I feel it's important to keep in mind that the hardware scheduler keeps dynamic statistics on port usage, so in a sense it's more like a JIT than a compiler, static analysis is only an approximation. What little experience I have told me it's always a good idea to compare IACA and LLVM-Mca's predicted schedule with a real profiler's output :) (Thanks for giving some nuance, it's appreciated. I'm not actually a compiler engineer or doing low-level magic for a living, so if you see anything wrong I would love to be corrected!)