3 ms·
Do you have a longer writeup on your experience performance tuning for the TPU? Is https://cloud.google.com/tpu/docs/performance-guide https://cloud.google.com/
by durst 5y ago
Do you have a longer writeup on your experience performance tuning for the TPU? Is https://cloud.google.com/tpu/docs/performance-guide https://cloud.google.com/tpu/docs/performance-guide the best resource in this area?
- sillysaurusx 5y agoSadly I have nothing to offer yet, other than to DM me on Twitter with questions. Always happy to answer basic ones; I love this stuff. But you’re right, I should write up something. That guide is pretty good, but it doesn’t walk you through anything specific; it shows you a map, but doesn’t take you on a trip, so to speak. Some tips: Use tf.name_scopes! You probably use them for variables, but there’s a different one for operations. If you make an @op_scope decorator, you can use it on all your ML functions and immediately get lots of insights as to where the XLA ops end up mapping to in your source code. As much as it pains me to say this, avoid tensorflow 2 style code like the plague. Pretend that if you use eager execution, someone will jump out of the bushes and shoot you. I technically use TF2.4 now, but it’s still session / graph-based, not the new tf.function magic. The new magic pipeline seems to be slower, harder to use, harder to debug, and very likely to explode when you do anything even slightly different than the tutorial examples. YMMV, and maybe things are better now (or in a future release). The profiler is magical; leverage it whenever you can. My workflow is to start up a TPU run, then ssh into my server and fire up a Tensorboard to that model dir. Then I manually navigate to <tensorboard_url>/#profile (because the “profile” button doesn’t seem to show up in the menu anymore) which then lets you “capture profile”. Make sure your TPU version matches your Tensorboard version. At this point we use TPU version 2.4.0, Tensorboard 2.4.1. If you get mysterious errors, this is likely the reason things are going wrong. Even TPU version 2.3.0 wasn’t enough. Happily, the new profiling tooling is totally badass. The trace viewer is great, the op profiling is ok (though I wish it would show me all the damn ops, instead of “helpfully” hiding all but the top N ops), and the memory viewer is incredible. You can see exactly where in your pipeline is causing “peak memory usage”, what the peak is (down to the kilobyte), and have at least some idea of what’s causing it. It’s not effortless though. The XLA fusion ops sort of make it harder to track down what’s doing what. (TF compiler is very powerful, but the trade off is that you almost never have manual control over memory usage, which can be frustrating). All in all (or all-to-all, ha) it’s a lot of fun if you like seeing expensive hardware go brrrr, as I do. It makes it all worth it when your loss drops from 11 to 3 overnight on a 430m parameter gpt model. :)
- durst 5y agoThank you for the tips! Whenever I hear about new accelerators, my first question is: "how do people in the real world run fast code on this"? Because (related to your above points), you don't use an accelerator for things to just run, they have to run fast. Otherwise, you'd just use a CPU. Detailed Question 1: do you have examples of fusion making things hard? Is there a way to nudge the compiler to not fuse or create a symbol table tracking fusion? Does fusion cause issues with the name_scopes? Detailed Question 2: isn't Pytorch eager execution? Do you know how it compares to Tensorflow's eager execution? General Request: I'm in a more theoretical position, writing papers on programming languages for accelerators https://aetherling.org/ https://aetherling.org/ and TAing courses on accelerators http://cs149.stanford.edu/fall20 http://cs149.stanford.edu/fall20. So, I'm excited to see people's practical experiences using these accelerators in industry. It would be very enlightening (if you have time) to write up a comparison of tuning a model for an A100's tensor cores vs a TPU. This seems like the key trade-off in comparing architectures.
- sillysaurusx 5y agoI'd be interested in contributing to the course, if you need anything specific. Here's a tensorboard URL that will probably stop working within a few days. http://bulma.tensorfork.com:31337/#profile http://bulma.tensorfork.com:31337/#profile You can view the memory profiler by using the dropdown menus on the left. Here's a particularly chonky CrossReplicaSum: https://i.imgur.com/CcdJzLj.png https://i.imgur.com/CcdJzLj.png The game here is to keep that number at the top -- peak memory usage -- below 15GB. In practice, TPUv3-8's run out of memory at around 14.5GB, which immediately crashes (and hence you can't profile it). So we're always trying to get as close to 15GB as possible. The first thing you immediately notice is that real-life training runs are very spiky. Different parts of the pipeline end up allocating wildly different amounts of memory. There's almost no such thing as a constant memory usage pipeline (which I was dismayed to discover). In this profiling run, you can see that there's a big ass-spike at ~4000 on the X axis. The green bar marks the lifetime of the operation causing the highest peak memory usage. Different operations depend on each other, forming a chain of allocations. a + b takes 'a' and 'b' as inputs, and any temporary tensor reachable by either 'a' or 'b' cannot be freed until a + b is finished executing. Ditto for all other operations. So you see, it's easy to accidentally build a "tower" of allocations, rather than a flat line. Thus, your total model parameter count is severely limited compared to what it could be, since in this situation the only way to reduce memory usage (without rewriting the code) is to scale down the model params. Hovering over the big ass-orange allocation, we see that the shape is -- gosh, tensorboard is infuriating sometimes. I tried to copy-paste the shape, but whenever I move the mouse off of the allocation, the info on the left vanishes. Anyway, the shape is F32[32,2048,1,12608][1,3,0,2]. It means the cross replica sum is happening across TPU cores 1, 3, 0, and 2; it's a float32 sum; the batch size is 32; the hidden dimension is 2048, and the vocab dimension is 12,608. Since it's across four cores, multiply that dim by 4, and the total vocab size is 50,432, which is exactly right for a GPT model (https://nv-adlr.github.io/MegatronLM https://nv-adlr.github.io/MegatronLM has details). So right away, we can see that (a) the non-peak memory usage is around 4GB or so, and (b) the peak mem usage of the spike is around 12GB. That means if we eliminate the spike, we can scale up our model by more than 3x, if usage scales linearly. (Sometimes you get lucky and it's linear, other times something is superlinear. It's more or less linear in my experience.) So how do we eliminate the spike? Heck if I know how the Google pros do it, but my way of doing it is to unstack along the batch dimension and perform each operation sequentially. In other words, the total memory usage here is O(32 * 2048 * 12,608) which is quite hefty. By unstacking along the batch dimension, you get 32 tensors, each of size 2048 by 12,608. Therefore, if you do each operation sequentially, the temporary buffer is now O(2048 * 12,608), giving us a 32x savings. Is this slower? Surprisingly, more often than not, it's as fast or faster. The reason is subtle: slowdowns occur due to memory bandwidth and network bandwidth. As long as the unstack is strictly a memory bandwidth effect, then it's just as fast, because you're trading CPU cycles for memory -- and you have tons of CPU cycles here, since it's a TPU core. (The TPU core utilization in our experience is always around 30%, and we've never seen it higher than 65%.) So you should always, always make this trade whenever possible. Network traffic is trickier. This is a cross-replica sum, which means it's sending the tensors across the network to each TPU core. The TPU cores are connected via a high speed interconnect nexus thingie, but like all bottlenecks, this one has a limit. It's a very high limit, but it's not endless. The only solution I've found is to think of an idea and then test that idea. Reasoning from first principles almost never works for me. I've seen others solve problems by reasoning from first principles, so it's possible that I'm simply stupid. But I find it's much more effective to try as many ideas as possible, as quickly as possible. You often end up surprised. I'll type more stuff later if I feel like it, or you can ask more followup questions. Feel free to DM me on twitter if you'd like to chat in realtime sometime.