5 ms·
Hey everyone, I'm the author of this article and I'm glad you've found it interesting. I'll be keeping an eye on this thread, so if you have any questions o
by RiotTony 11y ago
Hey everyone,
I'm the author of this article and I'm glad you've found it interesting. I'll be keeping an eye on this thread, so if you have any questions or comments I'll address them as soon as I can. I can already see some awesome questions here - looking forward to the discussion.
- Splines 11y agoWhat's the value gained from using vTune vs xperf sampled profiling? I use xperf and friends a lot and find that they're pretty good, but if vTune offers something substantially better I wouldn't mind taking a look at it. Also - interesting that you use the Chrome tools to visualize the graphs. I use WPA to view performance graphs and the breakdowns are somewhat similar. I think the regions of interest files can get you the rest of the way. Thanks for the writeup. It's interesting to see how other people tackle perf analysis.
- RiotTony 11y agoXperf or WPA is an excellent tool for holistically analysing your application (and all the other applications running on your machine at the same time). We do use that tool for looking at lots of different things: file IO, thread contention, server performance, etc. But for a single client running, I find VTune to be excellent. Its very well integrated into Visual Studio, and provides a number of different perf experiments that you can run to isolate the causes of your bottlenecks. VTune is commercial, but you can use Very Sleepy for a free alternative.
- josephg 11y agoWhy scan the table of precomputed values? It seems like the code would be both cleaner and faster like this: class AnimatedVariable { int numValues; std::vector<float> values; // ... } Then: if (!mPrecomputed) { float idx = time * numValues; float before = values[idx], after = values[idx+1]; return lerp(before, after, idx - (int)idx); } I guess there's a bit of float -> int coercion going on there, but it shouldn't be too bad.
- RiotTony 11y agoBecause we are interpolating between keys which aren't evenly distributed in time. What you have there is pretty much what we do when looking up the precomputed values from the table.
- deleted 11y ago[deleted]
- ralphael 11y agothanks for the article Tony. Loved it, really enjoyed the detail you went into. Look forward to more.
- sled 11y agoAre the data sources for Waffles and the Chrome visualizer different? It sounds like Waffles is real-time and the Chrome visualizer is not real-time. Do they use the same macros and route the buffer differently? Or do they have entirely different systems for gathering the profile data?
- RiotTony 11y agoWaffles is more than just a profiler. It is a nice, high level interface into our (non-public) debugging API. You are correct that the profiling info there is real-time, while Chrome is post. They all use the same buffers, but just interpret the data a little differently. There is also other profile info which is gathered by Waffles, stuff like number of visible particles, texture calls, GPU cost per emitter, amongst others, and that information is gathered through another interface.
- codesuki 11y agoNice article! I totally agree with the iterated approach to optimization (and refactoring). Since I am casually playing LoL I hope you will make it better and better ;) Talking about optimization, I always wonder why some people have those veeery long loading times. Actually I wonder what is loaded at all. Effects or something? Of course I have no idea of the insides of the game but looking from the outside the map and models should be easily cached. It would be great to hear more about those loading challenges!
- newman314 11y agoWhen you get a chance, look up the OODA loop. It's much similar to what you have written. When I first learnt about OODA, I was fascinated by how many areas I could suddenly see similar patterns. BTW the fellow behind OODA (Colonel Boyd) has a fascinating biography that's well worth a read.