3 ms·
TensorFlow Fold provides a TensorFlow implementation of the dynamic batching algorithm (described in detail in our paper [1]). Dynamic batching is an execution
by moshe 10y ago
TensorFlow Fold provides a TensorFlow implementation of the dynamic batching algorithm (described in detail in our paper [1]). Dynamic batching is an execution strategy for computation graphs, you could also implement it in PyTorch or Chainer or any other framework.
Our particular implementation of dynamic batching uses the TF while loop, which means that you don't need to make run-time modification to the actual TF computation graph. At runtime, we essentially encode the computation graph for (let's say) a parse tree as a serialized protocol buffer (tf.string), so instead of varying the computation graph itself we vary the input to a static computation graph instead. This particular implementation strategy is very much a byproduct of how TensorFlow works (static computation graph, heavy lifting happens in ops implemented in C++).
[1] DEEP LEARNING WITH DYNAMIC COMPUTATION GRAPHS, https://openreview.net/pdf?id=ryrGawqex https://openreview.net/pdf?id=ryrGawqex
- taliesinb 10y agoCongratulations on the nice work! It is very elegant to use combinators to formulate and solve this problem. Though my worry with combinators is that they introduce awkwardness with setting up 'long range' DAG edges, you can have to do a lot of shuttling of things around manually through tuples and whatnot, I'm not sure how it is with your framework. Am I right in thinking that there is no bucketing going on here? In other words, each batch is fixed length, and a short-lived DAG is planned and then simulated with tf.while to accommodate the set of shapes in that particular batch? Are there any problems when the input shapes are wildly different in expense? For example, imagine a size-agnostic convnet. Maybe some of the images in the training set are small, others are large, how would that look in your framework, if it can be done? Is junk padding part of picture, to help match almost-equal tensors to allow them to be batched?
- moshe 10y agoThanks, insightful questions. You're absolutely right about combinators and long-range dependencies. We have a special block type in the high-level API (https://github.com/tensorflow/fold/blob/master/tensorflow_fold/g3doc/py/td.md#td.metric https://github.com/tensorflow/fold/blob/master/tensorflow_fo...) for accumulating results here without having to explicitly shuttle them around; not perfect but very handy in many cases. Regarding your second question, the equivalent of padding in dynamic batching is "pass through" do-nothing ops, which are introduced transparently but worth being aware of if you want to understand the machinery. The worst case scenario here is chains of wildly varying lengths, where we need to add pass-throughs to the shorter chains to match the length of the longest chain.
- liuliu 10y agoFrom my understanding you cannot do the similar dynamic batching with PyTorch since it gets eagerly executed. Not sure about Chainer. Can you explain a bit more about how this can be applied to other frameworks?
- moshe 10y agoI've never used PyTorch or Chainer, so maybe I'm wildly wrong, but here goes.. Dynamic batching requires you to predefine a set of "batch-wise" operations (node types for computation graphs). This works fine for eager execution as well, because the eager code is inside a function definition. Of course if you want to do dynamic batching you need to introduce a layer of abstraction, e.g. instead of saying a+b you create a dynamic batching operation for addition and add edges to an execution graph, this is not different than in our TF implementation.
- liuliu 10y agoThanks. Yeah, my understanding is that you cannot have eager execution all the way down, some eager executions inside the function is fine, but you shouldn't have these outside the function blocks. Without this function block abstraction, it doesn't seem like you can implement dynamic batching very far though (or even at all).