2 ms·
great post. could you apply this same framework to optimize training as well?
by alanaan 3y ago
great post. could you apply this same framework to optimize training as well?
- varunshenoy 3y agoSlightly different set of trade-offs, but similar mental model. You always use large batch sizes (compute bound) and the bottleneck usually ends up communication between GPUs/nodes.