4 ms·
The memory constraint is certainly an issue, but I think they can be overcome with some software hacks (for a speed penalty, of course.) For example, gradient c
by bitforger 7y ago
The memory constraint is certainly an issue, but I think they can be overcome with some software hacks (for a speed penalty, of course.) For example, gradient checkpointing and gradient accumulation might help:
https://medium.com/tensorflow/fitting-larger-networks-into-memory-583e3c758ff9 https://medium.com/tensorflow/fitting-larger-networks-into-m...
https://medium.com/huggingface/training-larger-batches-practical-tips-on-1-gpu-multi-gpu-distributed-setups-ec88c3e51255 https://medium.com/huggingface/training-larger-batches-pract...