4 ms·
Saving memory using gradient-checkpointing
- supermdguy 9y agoI wonder if this would work when using keras with the tensorflow backend. Has anyone tried it?
- eggie5 9y agoIt takes about 4GB of memory to train a VGG network w/ batch size 8 on Tensorflow. Could I use a larger batch size at the expense of computation time w/ this module?
- iaw 9y agoYes, potentially a batch-size 64 with a 25% increase in run-time based on the figures they reported.