3 ms·
If I had to guess, based on other large models, it’s in the range of hundreds of GBs. It might even be in the TB range. To host that model for fast production
by binarymax 4y ago
If I had to guess, based on other large models, it’s in the range of hundreds of GBs. It might even be in the TB range. To host that model for fast production SaaS inference requires many GPUs. An A100 has 80GB, so a dozen A100s just to keep it in memory, and more if that doesn’t meet the request demand.
Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months.
- judge2020 4y ago> Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months. Based on this blog post where they scale to 7,500 'nodes', they say: > A large machine learning job spans many nodes and runs most efficiently when it has access to all of the hardware resources on each node. So I wouldn't be surprised if they do have a total of 7500+ GPUs to balance workloads between. TO add, OpenAI has a long history of getting unlimited access to Google's clusters of GPUs (nowadays they pay for it, though). When they were training 'OpenAI Five' to play Dota 2 at the highest level, they were using 256 P100 GPUs on GCP[0] and they casually threw 256 GPUs at 'clip' for a short while in January of 2021[1]. As for how they do it, see these posts: https://openai.com/blog/techniques-for-training-large-neural-networks/ https://openai.com/blog/techniques-for-training-large-neural... https://openai.com/blog/triton/ https://openai.com/blog/triton/ 0: https://openai.com/blog/openai-five/ https://openai.com/blog/openai-five/ 1: https://openai.com/blog/clip/ https://openai.com/blog/clip/