4 ms·
They use extensive model parallelism when training. Even TPUs (64GB) or Tesla V100 GPUs (32GB) don’t have enough memory to fit a model into a single child, so y
by jaredtn 6y ago
They use extensive model parallelism when training. Even TPUs (64GB) or Tesla V100 GPUs (32GB) don’t have enough memory to fit a model into a single child, so you’ll need activation checkpointing or model parallelism.