4 ms·
Is this the hardware used to build the model, or to run the model? Does this single 100000$ computer run the entire chatGPT service concurrently for all users a
by Valgrim 4y ago
Is this the hardware used to build the model, or to run the model? Does this single 100000$ computer run the entire chatGPT service concurrently for all users around the world? If so this seems incredibly cheap.
- paulluuk 4y agoIn machine learning, generally the expensive part is training the model, not running the model.
- bobleeswagger 4y agoAt this point the costs of hosting the trained model probably surpass the training cost.
- delijati 4y agohas anyone numbers on how many queries/second in inference time chatGPT can serve?
- paulluuk 4y agoIt might, I'm not sure. It depends on many factors: how much cloud discount are they getting, for example? And do you count the cost of training all previous models before chatGPT, or just the cost of training chatGPT itself? But you might be right, chatGPT is insanely popular right now.
- 0x008 4y agoThe magnificent part is that the 8xA100 is only a blade and you can fill a whole rack with them. Yet one is already incredibly powerful. Also the successor H100 is already released and at least 2x as powerful.
- kjs3 4y agoAlso the successor H100 is already released and at least 2x as powerful. Checks eBay for used A100s...will not be getting one soon. :-)
- HPsquared 4y agoSupercomputer used for training, top 5 in the world apparently: https://news.microsoft.com/source/features/ai/openai-azure-supercomputer/ https://news.microsoft.com/source/features/ai/openai-azure-s...
- anentropic 4y agoI would guess that's the minimum needed to run an 'instance' of the model for inference not sure how many concurrent users that could support, but I'd imagine the public service has a whole bunch of these racks from what I read... GPT-J at 6B params is too big to load on any consumer GPU (needs ~25GB) and then GPT-3 is 175B params or ~29x larger, so maybe needing 7-800GB RAM 8x 80GB A100s would give 640GB RAM, so its approx the right ballpark
- erichocean 4y ago> GPT-J at 6B params is too big to load on any consumer GPU (needs ~25GB) You can load that on M1/M2 Macs, which have unified GPU memory up to 128GB.