3 ms·
> just profile and find hot spots Do you mean "hot" in terms of CPU time? It may sound simple but in mature apps you don't have any obvious hot spots that you
by vient 3y ago
> just profile and find hot spots
Do you mean "hot" in terms of CPU time? It may sound simple but in mature apps you don't have any obvious hot spots that you can look at. Code is complex enough so just reading it also won't give you clues.
> Making it work well is simple, you just need to access memory linearly/contiguously ...
In real apps you do not usually have some kind of heavy linear array processing, instead they work with thousands of different objects depending on input - a smart prefetching can help but again, it is hard to find cache issues when you don't have obvious hot spots.
Perf can take stack traces on cache misses but the problem is that they are taken "asynchronously" so they do not point on exact instruction which caused cache miss - it's hard to analyze those records afterwards.
Hope I clarified what I meant: you can easily pinpoint slow functions in your program, with all sorts of debug info, but I don't know of a way to do the same efficiently for caching issues.
- CyberDildonics 3y agoIt may sound simple but in mature apps you don't have any obvious hot spots that you can look at. This is a defeatist attitude that is basically saying "it's impossible". I'm not sure what your expectations are, but people have been using profilers for a long time, so saying they magically don't work because your program is special is ridiculous. In real apps you do not usually have some kind of heavy linear array processing, instead they work with thousands of different objects depending on input This is again not true. If you have thousands of 'different objects' you should think about how you can replace them with a few different arrays of the data you are working on. This is not my idea, this is well worn performance advice. a smart prefetching can help It's not smart prefetching, it's just prefetching. If you access memory addresses next to each other sequentially, the prefetcher will grab memory ahead of the CPU and your program won't have to wait for it. it is hard to find cache issues when you don't have obvious hot spots. I don't think that's true at all, it just isn't important to deal with cache misses/pointer indirection unless it is repetitive. Memory access patterns are never this confusing to me, but I also plan for it ahead of time now that I have experience dealing with it. Perf can take stack traces on cache misses but the problem is that they are taken "asynchronously" so they do not point on exact instruction which caused cache miss - it's hard to analyze those records afterwards. You don't need an exact instruction and it probably wouldn't help you anyway. Memory access isn't a matter of a single instruction, it is about how the memory is layed out in the first place. you can easily pinpoint slow functions in your program, with all sorts of debug info, but I don't know of a way to do the same efficiently for caching issues. Once you weed out allocating memory in hot loops and slow IO, your slow functions and cache issues are probably the same thing. Also you are contradicting yourself here. You said: It may sound simple but in mature apps you don't have any obvious hot spots Then: you can easily pinpoint slow functions in your program If you have source code on github and you can tell me what lines are slow, I can probably tell you why.
- vient 3y ago> This is a defeatist attitude that is basically saying "it's impossible" I did not mean it like that. I am using profilers to find slow code successfully as well. The problem is that such profiler won't find you a spot where one object is slowly loaded from memory because CPU was not smart enough to prefetch it in advance, even if you profile with something as precise as Intel PT. > saying they magically don't work because your program is special is ridiculous I did not say that. Once again, I am talking about specific question of finding cache stalls in your program, not generic profiling. > If you access memory addresses next to each other sequentially It is of course beneficial to do it that way but it may be not trivial when you have complex relationships between different objects. Besides, the point of profiling is to efficiently find where you may need to do such optimizations, just saying "layout your objects optimally" does not help. > You don't need an exact instruction and it probably wouldn't help you anyway Exact instruction will help me to understand what specific object caused memory stall. Maybe this is indeed not that important. > Also you are contradicting yourself here I meant that you can easily find slow functions if there are any. If instead you have functiions a, b, c called sequentially and working same amount of time, you don't have obvious place to look at - if function b can be optimized by introducing cache awareness but other ones cannot, profile won't help you.
- CyberDildonics 3y agoI think the big picture here is that you are thinking it's difficult to notice cache misses and I think it isn't. Any time you are dereferencing pointers, working with heap allocated objects, vtables etc you can assume they are cache misses. If they are in a hot loop they could matter. Really any time you aren't running through contiguous memory or dealing with stack variables you should assume there are lots of cache misses. Your entire program is full of cache misses until you specifically structure them out.