3 ms·
So this looks like a further convergence of the tensorflow and pytorch APIs (the lower-level APIs at least). Tensorflow was designed with compilable graphs as
by ansk 4y ago
So this looks like a further convergence of the tensorflow and pytorch APIs (the lower-level APIs at least). Tensorflow was designed with compilable graphs as the primary execution model and as part of their 2.0 release, they redesigned the APIs to encompass eager execution as well. Pytorch is coming from the other end, with eager execution being the default and now emphasizing improved tools for graph compilation in their 2.0 release. The key differentiator going forward seems to be that tensorflow is using XLA as their compiler and pytorch is developing their own toolset for compilation. As someone who cares far more about performance than API ergonomics, the quality of the compiler is the main selling point for me and I'll gladly switch over to whatever framework is winning in the compiler race. Does anyone know of any benchmarks comparing the performance of pytorch's compilers with XLA?
- amelius 4y agoHow much performance can be squeezed from going from the plain python API to the graph-based solution, typically?
- ansk 4y agoThis varies quite a bit based on the type of model. The graph-based approach has two benefits: (1) removing overhead from executing python between operations and (2) enabling compilers to make optimizations based on the graph structure. The benefit from (1) is relatively modest for models which run a few large ops in series (e.g. image classifiers and most feedforward models) but can be significant for models with many ops that are smaller and not necessarily wired up sequentially (e.g. RNNs). In my experience, I've had RNN models run several times faster in tensorflow's graph mode than in its eager mode. The benefit from (2) is significant in almost any model since the typical "layer" building block (matmul/conv/einsum->bias->activation) can be fused together which improves throughput on GPUs. In my experience compilation can offer performance increases from 1.5x to 3x, but I don't know if this holds generally. Also note that the distinction between graph and eager execution can be somewhat blurry, as even an "eager" API could be calling a fused layer under the hood.
- PartiallyTyped 4y agoIt depends.. Using Jax to compile down to XLA, I often saw >2 orders of magnitude improvements. This however was roughly 6 months ago.
- brrrrrm 4y agoSome context/history: For compiler people reading this, a lot of common compiler terms have been entirely reinvented in the context of machine learning frameworks. An ML "graph" refers almost exactly to the dataflow graph (DFG) of a program. TensorFlow 1.0 only exposed a DFG, which is well known to be far simpler to apply optimizations to (assuming you have a linear algebra compiler). PyTorch integrated with Python (an interpreted language) and does not expose an underlying DFG. This is labeled "eager" and means that compilation of PyTorch requires optimization over both the control flow graph (CFG) and DFG. Python by default exposes neither of these things in a standard way. Some ML workloads simplify easily to a DFG (torch FX can handle this), but the general case does not. Although TorchScript (a subset of Python) tackled the CFG in 1.0, the team is now taking it further and compiling Python byte-code itself (with torchdynamo), which means you don't need to change any code and still get compilation speed ups! That's why 2.0 is significant. Of course, all of this requires a linear algebra compiler to actually do the optimizations which is why things like AITemplate (for inference) and TorchInductor (which calls into a bunch of other compilers for training) exist for PyTorch. TensorFlow's linear algebra compiler is XLA.
- brrrrrm 4y agoedit: TorchInductor calls into GCC/Triton depending. I mistook it for the various backends TorchDynamo supports (including TVM)