3 ms·
As someone from outside the AI space, how useful are chips like this from Amazon? Meaning am I going to have to rewrite my entire stack to make use of this and
by dvdbloc 3y ago
As someone from outside the AI space, how useful are chips like this from Amazon? Meaning am I going to have to rewrite my entire stack to make use of this and then be locked into Amazon forever? Or are workloads at least portable in some capacity?
- pram 3y agoThere’s semi lock-in. You can use torch, but you also have to use their framework. Not much different than using cuda I guess.
- belval 3y agoIt's a bit of a mixed bag. NeuronSDK is implemented with Torch XLA so there's not real lock-in there. Your code will usually just work on CUDA if it works in Trainium. However for larger models (ZeRO, Sharding, etc...) they have their own libraries that are more coupled to the actual accelerator topology because using something like DeepSpeed was not possible. Overall the challenge is often more with making your code work on their accelerator, not porting it away from their accelerators so I wouldn't be too concerned about lock-in.
- yeldarb 3y agoWe tried Inferentia & though they claim their compiler should "just work" on your existing models it was producing a bunch of NaNs for us & was opaque enough that we didn't really have a path forward so we gave up on it. If we were operating at 100x or 1000x the scale it might make sense to spend the extra engineering time, but we just pay for the NVIDIA chips (especially because a big chunk of our customers run at the edge where NVIDIA compatibility is even more important so we'd be doing double the engineering on an ongoing basis).
- bfeynman 3y agonot very useful, but it's getting better with XLA. It does not just "work" by default for any given model, and if you're missing an HLO that has an optimized kernel written you lose benefits of accelerated computing thus people won't use it. I would not put any effort into it until they truly have it where you just cast `to_device(..)`.