3 ms·
I don't but I'd like to know. I was under the impression that it was mostly GPU vram based but once the model is loaded, it could produce output quickly? I'm p
by smashed 3y ago
I don't but I'd like to know.
I was under the impression that it was mostly GPU vram based but once the model is loaded, it could produce output quickly? I'm probably over-simplifying things...
- soulofmischief 3y agogpt-3.5-turbo (default ChatGPT model) takes 8 A100s, ~$10k each. [0] The latest gpt-3.5-turbo model generates very quickly and cheaply (in part to some recently-discoverd optimization techniques... older versions cost 10x more). While the required hardware to run GPT-4 is currently unknown, it generates considerably slower on average and its much higher cost points to a higher hardware cost. And this is per request. It's bananas. [0] https://www.servethehome.com/chatgpt-hardware-a-look-at-8x-nvidia-a100-systems-powering-the-tool-openai-microsoft-azure-supermicro-inspur-asus-dell-gigabyte/ https://www.servethehome.com/chatgpt-hardware-a-look-at-8x-n...