4 ms·
One good starting point is the LLVM Machine Code Analyzer [1] What it does is use the scheduling info known to the LLVM optimizers to model how a particular CP
by vnorilo 5y ago
One good starting point is the LLVM Machine Code Analyzer [1]
What it does is use the scheduling info known to the LLVM optimizers to model how a particular CPU is going to execute your machine code.
I was lucky to start doing this back in the day when 486 was common and the Pentium was brand new. 80486 could do some instructions in parallel if you arranged them very carefully, and Pentium greatly boosted this capability. I used AMD CodeAnalyst (free) back then. I read Zen of Code Optimization by Mike Abrash which explained those particular microarchitectures very carefully. It may still be worth reading to understand how CPUs have evolved, but as this is uarch specific it will not be of great practical use.
The pipelines back then were simple enough to memorize so I spent some boring classes in senior high school plotting various software blitter algorithms on grid paper. Nowadays the superscalar capability is huge and you are better off taking a more statistical approach first - which execution units are stalled or underutilized - and see if you can tweak the instruction mix or find a false dependency that prevents register renaming.
For someone starting out I would recommend studying some smaller Arm chip that has limited superscalar capabilities. Sadly I can't name drop a book that would be a great help in that.
1: https://llvm.org/docs/CommandGuide/llvm-mca.html https://llvm.org/docs/CommandGuide/llvm-mca.html
- 10x-dev 5y agoThank you so much. This is awesome info. I'll check out LLVM MCA and I've also ordered the book. Happy holidays!