3 ms·
The reason all you hear is "profile profile profile" is that the performance implications of a given code change may sometimes be completely non-obvious, due to
by ehsanu1 12y ago
The reason all you hear is "profile profile profile" is that the performance implications of a given code change may sometimes be completely non-obvious, due to all the complex trickery in the processor and cache hierarchy. So what do you do after profiling and finding a performance bottleneck? If you can't identify exactly what the bottleneck's cause is and the solution is not obvious, I suppose all that's left is to do a best guess of the cause and think of a possible solution. Then profile after implementing your solution to check if it made a (positive) difference. And do the profiling/benchmarks even if the solution seems obvious, since it's so easy to be wrong.
Detecting cache misses can be done with valgrind: http://valgrind.org/docs/manual/cg-manual.html http://valgrind.org/docs/manual/cg-manual.html
There are architecture-specific instruction-level profilers as well. This video mentions a couple and is also a great overview of how non-intuitive performance can be with modern CPUs: http://channel9.msdn.com/Events/Build/2014/4-587 http://channel9.msdn.com/Events/Build/2014/4-587