4 ms·
Most of the time I've written performance sensitive C or C++, I spend most of the time looking at disassembly or pipeline models, trying to spin up an instructi
by vnorilo 5y ago
Most of the time I've written performance sensitive C or C++, I spend most of the time looking at disassembly or pipeline models, trying to spin up an instruction mix that will utilize the CPU backend in a good way.
In the best case, the compiler finds the instructions I want, and often it does a very good job with details like register allocation.
So the abstraction leaks and inverts like hell. But still, in many cases the C program can still be portable, even if none of the performance tuning is - it targets the wrong pipeline and memory subsystem traits.
When you need to rely on intrinsics, it is different - but I'd say intrinsics are closer to inline asm than C. Saying that as someone who might have ported a DSP library from SSE4 to AltiVec with mostly #define. I know. Also I'm happy it wasn't the other way around!
- 10x-dev 5y agoSlightly off topic question: how can I get started with pipeline models and with optimizing the instructions that get executed? What tools would one use to inspect what gets executed, trace the instructions and measure the execution time? Could you recommend any books on this topic? Thanks!
- vnorilo 5y agoOne good starting point is the LLVM Machine Code Analyzer [1] What it does is use the scheduling info known to the LLVM optimizers to model how a particular CPU is going to execute your machine code. I was lucky to start doing this back in the day when 486 was common and the Pentium was brand new. 80486 could do some instructions in parallel if you arranged them very carefully, and Pentium greatly boosted this capability. I used AMD CodeAnalyst (free) back then. I read Zen of Code Optimization by Mike Abrash which explained those particular microarchitectures very carefully. It may still be worth reading to understand how CPUs have evolved, but as this is uarch specific it will not be of great practical use. The pipelines back then were simple enough to memorize so I spent some boring classes in senior high school plotting various software blitter algorithms on grid paper. Nowadays the superscalar capability is huge and you are better off taking a more statistical approach first - which execution units are stalled or underutilized - and see if you can tweak the instruction mix or find a false dependency that prevents register renaming. For someone starting out I would recommend studying some smaller Arm chip that has limited superscalar capabilities. Sadly I can't name drop a book that would be a great help in that. 1: https://llvm.org/docs/CommandGuide/llvm-mca.html https://llvm.org/docs/CommandGuide/llvm-mca.html
- 10x-dev 5y agoThank you so much. This is awesome info. I'll check out LLVM MCA and I've also ordered the book. Happy holidays!
- vnorilo 5y agoAh, one resource that slipped my mind was Jon Stokes' microarchitecture articles [0] at Ars Technica when it was still good (it's all gadget/lifestyle/policy stuff nowadays). Jon also has this book which I seem to remember was fairly good [1]. Don't be put off by the age - uarch on CPU side has mostly been more and more of the same for x86 chips. 0: https://arstechnica.com/author/hannibal/ https://arstechnica.com/author/hannibal/ 1: https://www.goodreads.com/book/show/610830.Inside_the_Machine https://www.goodreads.com/book/show/610830.Inside_the_Machin...
- lapinot 5y agocompiler inspector is fun: https://gcc.godbolt.org/ https://gcc.godbolt.org/ perf related tools on linux: https://perf.wiki.kernel.org/index.php/Main_Page https://perf.wiki.kernel.org/index.php/Main_Page Didn't read it but it's a well known reference: the dragon book (Compilers: Principles, Techniques, and Tools) 2nd edition has stuff on machine-dependent optimizations (chapters 10 and 11 apparently). Most likely a good read.
- astrange 5y agoThe dragon book is a bad compiler book. CPU pipelines are extremely dynamic since the slowest actions affecting them (memory reads) are also the least predictable (depends on what ends up in the cache). So it’s usually not worth trying to control them so precisely, but knowing how the first layers work can be good. For x86 the best resources are Agner Fog’s manuals and then the official Intel/AMD ones.
- vnorilo 5y agoAgreed with "usually". When you model pipelines, you must know the cache behaviour of your workload, otherwise it is a waste of time. But when you do, in order to approach theoretical machine limits, you do need airtight core resource utilization in your kinner loops.
- optymizer 5y agoContext for others who may have not read "Compilers" by Aho: it's a great textbook and a wonderful resource to learn about compilers, but I wouldn't call it a "good read". A good read is "The Pragmatic Programmer", for some. The dragon book is raw knowledge and it's filled with proofs. The content in the dragon book is the equivalent of a two semester course on compilers, and that's with a prof and TA. Ideally, to get the full benefit of reading this book, you need to reserve a year of your free time after 5pm and be ready to build an optimizing compiler. It's one of those books that requires your full attention for an extended period of time and you come out the other side a stronger developer, only because it didn't kill you with knowledge. Edit: I realized now you may be saying the 2 chapters (10 and 11) are most likely a good read to learn about optimizations, not that the entire book is a "good read" in general. Makes sense - I'll leave the comment up, with the disclaimer that I'm referring to reading the entire book cover to cover.
- gavinray 5y agoWhat is a "pipeline model"? Google doesn't give me any relevant results EDIT: Think I found an answer > "Some very experienced programmer from another company told me about some low-level code-optimization tips that targeting specific CPU, including pipeline-optimization, which means, arrange the code (inlined assembly, obviously) in special orders such that it fit the pipeline better for the targeting hardware." From https://stackoverflow.com/questions/14657247/pipeline-optimzation-is-there-any-point-to-do-this https://stackoverflow.com/questions/14657247/pipeline-optimz... Answer in comments point to this as a resource: https://www.agner.org/optimize/ https://www.agner.org/optimize/
- capitalsigma 5y agoGP probably means something like IACA
- vnorilo 5y agoWhat I think of as "pipeline model" is similar to what LLVM-MCA produces in its "Timeline View" [1] It basically tries to model statically how instructions travel through the pipeline of a particular CPU, which is useful for finding bottlenecks. 1: https://llvm.org/docs/CommandGuide/llvm-mca.html#timeline-view https://llvm.org/docs/CommandGuide/llvm-mca.html#timeline-vi...
- Koshkin 5y ago> trying to spin up an instruction mix But this may change from compiler to compiler and from version to version of the same compiler. I think you may be better off writing assembly code directly rather than trying to coerce the compiler to do exactly what you want.
- vnorilo 5y agoYou'd be surprised how stable it actually is in practice. The occasional big swings in benchmarks tend to be due to compiler A pattern-matching an idiom that compiler B does not. Regressions from A.1 to A.2 are rare, and usually either bugs or that optimizer default target has shifted and now neglects whatever uarch you regressed on.