4 ms·
Realistically you'd be able to train up to 760M param models. For that you'd need 32gb+ VRAM GPUs which I think AWS might have. You can try iwth 16gb VRAM GPUs,
by sdan 6y ago
Realistically you'd be able to train up to 760M param models. For that you'd need 32gb+ VRAM GPUs which I think AWS might have. You can try iwth 16gb VRAM GPUs, but you would need to figure out FP16.
https://github.com/shawwn https://github.com/shawwn is doing some work in the GPT-2 space including using TPUs instead -- which has given him pretty good results.
- jaredtn 6y ago32GB RAM is nowhere near what's necessary. It can barely fit the GPT-2 models on there with a batch size in the single digits. We'll need extensive model parallelism libraries (Zero2 from Microsoft) to run this at all.