3 ms·
Serving a request to one of their APIs requires orders of magnitude more compute than your typical web service
by dontreact 4y ago
Serving a request to one of their APIs requires orders of magnitude more compute than your typical web service
- la64710 4y agoReally ? How do you know? Have they shared any credible data around it? Based on experience of BERT , yes maybe to get the best experience or to serve millions of users you need to run any model on compute intensive infrastructure , BUT if you just want to run for yourself and do some small testing you can very well download it from huggingface and elsewhere and run it on your laptop.
- Rebelgecko 4y agoI've tried a few flavors of GPT on my laptop and a single request used orders of magnitude more CPU and RAM than running a "SELECT * FROM WHEREVER" SQL query
- TOMDM 4y agoPretty much any laptop would take hours to get anything significant done with GPT3 as the model would need to be batched in and out of memory from disk. The amount of memory required to run these models is immense. If you want a comparison, try running the largest version of the open source BLOOM model yourself.
- localhost 4y agoGPT3 is ~175B parameters. At float16 precision, that's 350GB of weights. BLOOM-176B is about the same size. Here's one person's experience; a token is ~0.75 words. "The Python code in this tutorial generates one token every 3 minutes on a computer with an i5 11gen processor, 16GB of RAM, and a Samsung 980 PRO NVME..." [1] https://towardsdatascience.com/run-bloom-the-largest-open-access-ai-model-on-your-desktop-computer-f48e1e2a9a32 https://towardsdatascience.com/run-bloom-the-largest-open-ac...
- Waterluvian 4y agoGoodness... And to think my brain was free!
- tstrimple 4y agoI wonder if your parents would agree with this statement.
- Waterluvian 4y agoI just asked. My dad said, “more or less. But your older brothers were very expensive.”