4 ms·
How would one learn about this depth of C and its optimizations? I'm interested in it as an art form, something to do in my spare time.
by YesBox 4y ago
How would one learn about this depth of C and its optimizations? I'm interested in it as an art form, something to do in my spare time.
- bfz 4y agoapt-get install linux-tools perf record -g myapp perf report Play with the output for a few weeks. "Why is this slow?" is a never-ending question leading down infinite corridors and weird specialist tools. Cachegrind and VTune are also instructive to play with
- joshspankit 4y agoStart with assembly
- Maursault 4y agoMachine code is only moderately less convenient.
- panic 4y agoFamiliarize yourself with the tools for measuring performance on whatever platform you want to optimize for. Then you can just start messing around and seeing what kinds of things make the number go down.
- nomel 4y ago> Then you can just start messing around and seeing what kinds of things make the number go down. And this will almost certainly require a deep understanding of the instruction set, and dropping into inline ASM, periodically.
- panic 4y agoOptimizing memory layout and reducing allocations are pretty ISA-independent. But yeah, at some point you will have to start looking at the assembly to optimize further (even if only to know when to tell the compiler to stop inlining a function that causes tons of register spilling inside a tight loop).
- saagarjha 4y agoIn most cases there’s a lot you can do by just peeking at the source code. Almost all programs ship without having been run through a profiler at all.
- emsy 4y agoC is not that deep. In the end its just something that produces assembly/machine code: jumps, reads writes etc. you need to get a feeling for what the compiler outputs for a given input (try Godbolt!). Then you need to know how those outputs will perform on your target hardware. This video is a good example: it states, modern compilers unroll loops, because they usually perform better on modern systems, but not the N64. So knowing your hardware is just as important.
- MaxBarraclough 4y ago> C is not that deep. It absolutely is. Look at its memory model, or look at any StackOverflow question or HackerNews discussion about edge-cases of undefined behaviour. Example: [0]. Alternatively, look at challenging interview questions that require a deep understanding of the C language. C gives the appearance of being a simple language, but has many dark corners that many programmers are unaware of. [0] https://news.ycombinator.com/item?id=22867059 https://news.ycombinator.com/item?id=22867059
- saagarjha 4y agoVery little of this is relevant for performance. It’s important to know what’s UB so the compiler doesn’t miscompile your code, but most performance improvements don’t come from “oh the compiler can run aliasing analysis on this better so it’s 10x faster” but “this loop is O(n^3)” or “I should convert this linked list to a flat array”.
- emsy 4y agoThat's basically what I meant. But even taking all the ugly parts of C into consideration, they are shallow. You don't need to dig into 10 layers of abstraction to get to the bottom of things. And the example posted is about what others say about how to interpret the C standard. I don't care. At the end of the day you can look at the asm output and understand what's happening for the compiler you're using.
- ninjinxo 4y agoRead Agner Fog's optimisation manuals. https://www.agner.org/optimize/#manuals https://www.agner.org/optimize/#manuals
- klodolph 4y agoA big part of it is learning computer architecture. If you have a solid grasp of things like how memory access works and how the CPU pipeline works, you can start to get a better mental picture of how fast a particular piece of code will run based on what instructions the CPU is executing and what the memory layout is. There are textbooks and classes on computer architecture. Funny enough, many of them use MIPS, which is the architecture used in the N64. Optimizing for N64 also requires understanding how the RCP works, which is a separate topic.
- djmips 4y agoOld skool book recommendation. Not necessarily C, more ASM but still a good book for the early nineties optimization that still has a mentality that's still valid. I believe that people today would be more concerned about how fast their code runs if it was visible, that is, if they profiled it with good tools! https://www.amazon.com/Inner-Loops-Sourcebook-Software-Development/dp/0201479605 https://www.amazon.com/Inner-Loops-Sourcebook-Software-Devel...
- fps_doug 4y agoSome of that doesn't have to do with C even. Like, changing memory layout of variables, or the way you access data. A simple example is if you have a 2D array, and you have the data from the individual rows consecutively in memory, but then you loop over the columns in your outer for-loop and over the rows in your inner for-loop. This means you access the first element from the first row, then the first element from the second row, then the first element from the third row, and so on. All these elements are far apart in memory, but every time you access one element, let's assume it's an uint32_t, the CPU fetches a whole cache lane of e.g. 64 bytes and puts it in the CPU cache in anticipation that you access data close to this in the near future. But you don't, so the CPU has to fetch another 64 bytes block for the first element of the second row, uses only 4 bytes from that, and so on. If your 2D array is large enough, by the time you finish the first iteration of the inner loop and start reading the second element of every row, the 64 byte cache lane that was fetched when you read the first element of the first row has already been evicted from the CPU cache again when you read the first element of row 2000, so the same 64 byte block has to be fetched from RAM again. This makes a huge performance difference, and is applicable to pretty much every programming language.