4 ms·
We currently have google engineers training on TPU pods with PyTorch Lightning. TPU support is VERY real... but yes, sometimes it breaks but PyTorch and Google
by wfalcon 6y ago
We currently have google engineers training on TPU pods with PyTorch Lightning.
TPU support is VERY real... but yes, sometimes it breaks but PyTorch and Google are working very hard to bridge that gap.
But we have dedicated partners at Google on the TPU team working to get Lightning working seamlessly on pods.
Check out the discussions here:
https://github.com/PyTorchLightning/pytorch-lightning/issues/3058 https://github.com/PyTorchLightning/pytorch-lightning/issues...
- sillysaurusx 6y agoNo, you do not support the TPU infeed, and this is a crucial distinction. Saying that you do support this has caused endless confusion and much surprise. It’s almost not an exaggeration to say that you’re lying (sorry for phrasing this so bluntly, but I’ve seriously spent dozens of hours trying to break this misconception due to hype like this). TPU support is real. Pytorch does in fact run on TPUs. But you don’t support TPU CPU memory, the staging area that you’re supposed to fill with training data. That staging area is why a TPU v3-512 pod can train an imagenet resnet classifier in 3 minutes at around 1M examples per second. You will not get anywhere near that performance with pytorch on TPUs. In fact, you’re expected to create a separate VM for every 8 TPU cores. The VMs are in charge of feeding the cores. That’s insane; I’ve driven TPU pods from a single n1-standard-2 using tensorflow. Repeat after me: if you are required to create more than one VM, you do not (yet!) support TPU pods. I wish I could triple underline this and put it in bold. People need to understand the limitations of this technique. Creating 256 VMs to feed a v3-2048 is not sustainable.
- minimaxir 6y agoThere's a difference between "supporting TPUs" and "supporting TPUs at 100% potential". Although the distinction is important, I don't think the marketing here is misleading.
- sillysaurusx 6y agoNot only is it misleading, it even somehow tricked you. :) We’re not talking about a small 10% reduction in performance here. We’re talking like 40x differences. If it seems unbelievable, and like it can’t possibly be true, well: now you understand my frustration here, and why I’m trying to break the myth. Notice not a single benchmark has ever gone head to head in MLPerf using pytorch on TPUs. And that’s because using pytorch on TPUs requires you to feed each image manually to the TPU on demand, from your VM. Meaning the TPU is always infeed bound. Engineers should be wincing at the sound of that. Especially anyone with graphics experience. Being infeed bound means you have lots of horsepower sitting around doing nothing. And that’s exactly the situation you’ll end up in with this technique. There’s a way to settle this decisively: train a resnet classifier on imagenet, as quickly as possible. If you get anywhere near the MLPerf v0.6 benchmarks for tensorflow on TPUs, I will instantly pivot the other direction and sing the praises of pytorch on TPUs far and wide.
- deleted 6y ago[deleted]
- wfalcon 6y agoLike I said... pytorch and tensorflow team are working very hard to make this work. And yes, it's not a 1:1 with tensorflow, but we're making progress very aggressively.
- sillysaurusx 6y agoI love what you guys are doing, and I love improving the ML ecosystem, but you’ve godda understand, people see this and think “oh, ok, it’s a small difference, no big deal.” In fact it’s a huge difference. Picture a person with one arm and without legs. Would you say they aren’t “1:1 in terms of features”? They certainly won’t be winning any races. And unlike real people, you can’t graft on a prosthetic limb to help this situation. The issue I’m describing here is a fundamental one that everyone keeps trying to sweep under the rug and pretend isn’t an issue. And then everyone wonders what’s going on.
- wfalcon 6y agoI 100% agree. We don't want to misrepresent TPU support. In fact, we explicitly warn users in our docs. Open to suggestions about how we can communicate this much better to our users. We just need to be a part of the effort to help bridge the big gap and barriers keeping users from TPU adoption. https://pytorch-lightning.readthedocs.io/en/latest/tpu.html#tpu-support https://pytorch-lightning.readthedocs.io/en/latest/tpu.html#...
- judge2020 6y ago> In fact, we explicitly warn users in our docs This was mentioned above, but nowhere on that page does it talk about any limitations whatsoever.