3 ms·
Latency in training literally does not matter. You care about throughput. In serving, where latency matters, most DL frameworks allow you to serve the model fro
by kahnjw 7y ago
Latency in training literally does not matter. You care about throughput. In serving, where latency matters, most DL frameworks allow you to serve the model from a highly optimized C++ binary, no python needed.
The poster you are replying to is 100% correct.