2 ms·
how big do the models get? are you able to say?
by robmsmt 6y ago
how big do the models get? are you able to say?
- boulos 6y agoAs big as you want, kind of. The challenge, as in large-scale physics, is how many nodes you can stick together with sufficiently high bandwidth (low latency is less important in the ML space, because there are lots of ops per byte, unlike some CFD Simulations that have very few per update). On our Cloud TPU product page [1], we have a single TPU v3 pod with 32 TB of memory. For the most recent MLperf submission, the TPU folks hooked up four of them [2]. There’s obviously a reduction in scalability from doing so (see weak scaling versus strong scaling terminology), but that’s the interesting co-design question: what kind of models can you usefully train in an “even more distributed” mode? Outside of TPUs though, even our single 16x A100 offering has 640 GB all connected by NVLINK (other providers went with 8, so 320 GB of “system memory”) and there are at least a few in a single rack. So the era of TiB scale models is certainly “semi feasible” and “open to all”. The challenge is that you need to also train these for quite some time. 1000 V100s would cost you at least $2000/hr to rent. Many models are sufficiently complicated (not just large) that you end up training them for days and weeks, even with this much compute. So the numbers add up quickly. But just being “big” doesn’t mean “trained for a month on a supercomputer”. [1] https://cloud.google.com/tpu https://cloud.google.com/tpu [2] https://www.google.com/amp/s/cloudblog.withgoogle.com/products/ai-machine-learning/google-breaks-ai-performance-records-in-mlperf-with-worlds-fastest-training-supercomputer/amp/ https://www.google.com/amp/s/cloudblog.withgoogle.com/produc...
- divtiwari 6y agoHi, where can I learn more about Distributed Systems? I guess your job requires that knowledge. Any accessible resources for a fresh grad?