3 ms·
> and this benchmark seems to back his claims. These benchmarks evaluate single-node performance. LeCun's remarks were concerning distributed training (specifi
by rryan 11y ago
> and this benchmark seems to back his claims.
These benchmarks evaluate single-node performance. LeCun's remarks were concerning distributed training (specifically that bandwidth between machines is a limiting factor to scalability) -- which we can't test yet since the current version of TF is single-node only.
Dean's response in the video "it depends on your [computer] network" is an interesting response :).
- kastnerkyle 11y agoBoth Google and FB (from what I understand) have a ton of tricks to help this, but I expect Yann was speaking generally even with all these tricks, and he is right from what I have seen. The old paper on DistBelief talks about topping out at ~80 machines due to network overhead - it would be great if they talk more about distributed TF in an upcoming paper. If TF really has a way to make general, networked, distributed training efficient (more than 1 bit weight updates, low precision weights and all the other crazy tricks which already exist) - that is truly remarkable and they rightly deserve huge kudos. If they are faster in distributed training only on Google machines or with Google's network architecture, that isn't really a useful datapoint for the general public. We could wire everything with 10GB/s NiCs, change to jumbo frames, pull all these other bandwidth reducing tricks, trick out the Linux kernel, etc. and then your network probably won't matter - but that isn't really general or cheap. The key will be what is the minimum effort necessary to avoid network bottlenecking, and does TF improve that minimum level over existing solutions?