3 ms·
This varies quite a bit based on the type of model. The graph-based approach has two benefits: (1) removing overhead from executing python between operations a
by ansk 4y ago
This varies quite a bit based on the type of model. The graph-based approach has two benefits: (1) removing overhead from executing python between operations and (2) enabling compilers to make optimizations based on the graph structure. The benefit from (1) is relatively modest for models which run a few large ops in series (e.g. image classifiers and most feedforward models) but can be significant for models with many ops that are smaller and not necessarily wired up sequentially (e.g. RNNs). In my experience, I've had RNN models run several times faster in tensorflow's graph mode than in its eager mode. The benefit from (2) is significant in almost any model since the typical "layer" building block (matmul/conv/einsum->bias->activation) can be fused together which improves throughput on GPUs. In my experience compilation can offer performance increases from 1.5x to 3x, but I don't know if this holds generally. Also note that the distinction between graph and eager execution can be somewhat blurry, as even an "eager" API could be calling a fused layer under the hood.