3 ms·
For ConvNets, the memory use of the models themselves is pretty modest. For example, even with 0.5B parameters, with FP32, weights+gradients+momentum should use
by trott 6y ago
For ConvNets, the memory use of the models themselves is pretty modest. For example, even with 0.5B parameters, with FP32, weights+gradients+momentum should use just 6GB (unless your framework sucks, or you have extra overhead from distributed training) So, if your model is twice smaller, you'll only save 3GB. If your VRAM is 32GB, saving 3GB won't let you use a much bigger batch size. On the other hand, the absence of batch norm can actually lead to memory savings proportional to batch size.