4 ms·
Using any current architecture it is infeasible to do backprop (training) due to the massive communication requirements. Inference is possible to do in sharded
by f_devd 4y ago
Using any current architecture it is infeasible to do backprop (training) due to the massive communication requirements. Inference is possible to do in sharded way but still not as practical as just loading the model weight that are needed on-demand from disk; still a distributed job queue being processed may be beneficial depending on the throughput/costs required by researchers.