4 ms·
not very useful, but it's getting better with XLA. It does not just "work" by default for any given model, and if you're missing an HLO that has an optimized k
by bfeynman 3y ago
not very useful, but it's getting better with XLA. It does not just "work" by default for any given model, and if you're missing an HLO that has an optimized kernel written you lose benefits of accelerated computing thus people won't use it. I would not put any effort into it until they truly have it where you just cast `to_device(..)`.