3 ms·
The use of gradient accumulation here is interesting, and I'd argue maybe not optimal. He uses a batch size of 8 (because his GPU only has enough RAM to allow
by jphoward 6y ago
The use of gradient accumulation here is interesting, and I'd argue maybe not optimal.
He uses a batch size of 8 (because his GPU only has enough RAM to allow 8 CT scans to be tracked through the network at once). He acknowledges that this is smaller than he would like, and so he uses gradient accumulation so that he only adjusts the weights of the network every 8th forward pass. Effectively, he averages the results over 8 batches.
The thing about batch size is it's a trade off. It turns out, even if your GPU could support you putting your entire dataset through at once, this isn't a good idea. The explanation for this can be simplified as "if you give it the same data every time, it'll give you the same answer every time". This may sound fine, but actually a bit of randomness is very helpful to get the network outs of local minima, i.e. local ruts. This is why we do "mini batch" stochastic gradient descent. Now his dataset is only 200something patients and 300something scans, meaning he only gets 4-6 batches out of his dataset for every epoch (every cycle). I would expect (though I have no proof) that using a lower batch size might increase stochasticity here.
The problem is, smaller batch sizes also have their problems. The main reason for this is due to the ubiquity of a layer in CNNs called batch normalisation (BatchNorm) layers. These layers calculate on the fly (during training) what the mean and standard deviation across every feature in the current batch is, and rescale them so together they mean of 0 and a standard deviation of 1. The problem is, when your batch size is very small, your means and standard deviations might be wildly inaccurate (e.g. if your batch size is 4, it wouldn't be that surprising if every sample in the batch is a COVID patient, or every sample is normal). The batchnorm layers will 'remap' those images onto a normal distribution, and make a batch of 4 COVID patients look more normal, and nice versa. For this reason, we try not to train networks with batch sizes below 10, because when BatchNorm is involved, things tend to start behaving badly. Unfortunately, BatchNorm doesn't give a damn about gradient accumulation, because they are still different batches. Gradient accumulation does not get around the biggest problem of small batch sizes, and in my experience very rarely helps. Instead, you're better of using a different normalisation layer.
- jcreinhold 6y agoRegarding smaller batch sizes and batch normalization: Have you found another normalization layer to work better for small batch sizes? I agree that the mean and variance of the small batches won't be representative of the true mean and variance, but in practice, I've used batch norm for small batch sizes successfully (e.g., <8). In medical imaging, due to memory constraints, I commonly see batch sizes of 2 (or even 1, although it's not really "batch norm" at that point). The paper "Revisiting small batch training for deep neural networks" [1] discusses the benefits of small batch sizes even in the presence of batch norm (see Fig. 13, 14). They only look at some standard CV datasets, so it isn't conclusive by any means, but the experimental results jive with my experience and what appears to be other researchers experience. [1] https://arxiv.org/pdf/1804.07612.pdf https://arxiv.org/pdf/1804.07612.pdf