7 ms·
This is cool! But I wonder if it's economical using cloud hardware. The author claims 1.12 tokens/second on the 175B parameter model (arguably comparable to GPT
by muttled 4y ago
This is cool! But I wonder if it's economical using cloud hardware. The author claims 1.12 tokens/second on the 175B parameter model (arguably comparable to GPT-3 Davinci). That's about 100k tokens a day on the GCP machine the author used. Someone double check my numbers here, but given the Davinci base cost of $0.02 per 1k tokens and GCP cost for the hardware listed "NVIIDA T4 (16GB) instance on GCP with 208GB of DRAM and 1.5TB of SSD" coming up to about $434 on spot instance pricing, you could simply use the OpenAI API and generate about 723k tokens a day for the same price as running the spot instance (which could go offline at any point due to it being a spot instance).
Running the fine-tuned versions of OpenAI models are approximately 6x more expensive per token. If you were running a fine-tuned model on local commodity hardware, the economies would start to tilt in favor of doing something like this if the load was predictable and relatively constant.
- swatcoder 4y agoSometimes control is more important than cost.
- cypress66 4y agoThis is most likely aimed at people running models locally. And a homelab with 3090s/4090s is one or two orders of magnitude cheaper than GCP, if you use them continuously.
- SomeHacker44 4y agoI do not know anyone offhand with a 200+GB RAM home computer. The GPU is not all that is needed; you need to keep the parameters and other stuff in memory too.
- Filligree 4y agoRunning it off a fast NVMe apparently works. I don't know what the performance is like, though.
- zargon 4y ago256gb of ddr4 rdimms only costs about $400 right now. $200 for ddr3. Not uncommon in homelabs. I don't think 200gb ram is actually required, that's just what that cloud vm was spec'd with. Though the 175b model should see benefit with ram even beyond 200gb.
- woadwarrior01 4y agoI own a two year old headless home computer with 256GB of RAM and two 3090s. I ssh into it from my mac to run ML training jobs.
- hedgehog0 4y agoWhat’s the price may I ask
- alex_sf 4y agoI have a similar setup. GPUs were $1.6k US, rest of the system was another ~$1k. This includes 256GB RAM and dual Xeons.
- zac_hudson 4y agoI bought an old server from ebay, and ripped the guts out to make a desktop computer. 128GB of ddr3 ram and 48 E5 v2 xeons for 300$.
- pclmulqdq 4y agoCloud accelerators carry a huge price premium because there aren't very many of them available and they aren't as fungible as CPUs. Comparing to a local GPU would likely be favorable for the local machine.
- breckenedge 4y agoThanks for running the cloud numbers on this. I ran some DIY numbers and they indicate less than a week to break even with the cloud, including all hardware and electricity costs. The cloud seems stupid expensive compared to running your own hardware for this kind of task.
- throwawayapples 4y agoThe cloud is always stupid expensive compared to running your own hardware for almost any sort of task that isn't highly variable upon one or more axis (cpu, ram, etc), but less than a week to break even is truly shocking.
- p1esk 4y agoThe cloud has been cheaper to train large models on for a couple years now. Compare buying 8xA100 server vs renting one on Lambda Labs. At least 3 years to break even - if you are using it non-stop 24/7. Longer if not.
- cardine 4y agoThis is not true - the break even period is closer to 6-7 months.
- p1esk 4y agoA single 8xA100 server is ~150k. On demand cost to rent it is $8.8/hour. Do the math and don't forget the energy costs.
- cardine 4y agoI'd suggest finding a cheaper vendor if that is the lowest price you can get for an 8xA100 server. We spend a lot on both and colo our servers so I've definitely done the math!
- 4y ago
- ImprobableTruth 4y agoYou've made one huge mistake: Davinci's $0.02 is not just per 1k tokens generated but also context tokens consumed. So if you generate 50 tokens per request with 1k context, the price is actually 20 times as large at $0.40 per 1k tokens generated - much less palatable, costing 3 times as much as the cloud hosted version of this. And that's not even taking into account the gigantic markup cloud services have.
- yorwba 4y agoMost of the computational cost of producing an output token is spent on consuming input tokens (including previous output tokens that are fed back in); only the final unembedding matrix could be eliminated if you don't care about the output logits for the context. So it's not correct to only modify OpenAI's prices to account for the ratio of context tokens to output tokens. Both of them get multiplied by 20 (if that's what your ratio is).
- ImprobableTruth 4y agoNo, because they're already taking that into account. >Metric: generation throughput (token/s) = number of the generated tokens / (time for processing prompts + time for generation). (Though they're doing batching, so this is an unfair comparison. Would be interesting to get single batch speed.)
- borzunov 4y agoI'm afraid that, unlike proprietary APIs and Petals, this system can't be used for single-batch inference of 175B models with interactive speeds - the thing you actually need for running ChatGPT and other interactive LM apps. See https://news.ycombinator.com/item?id=34874976 https://news.ycombinator.com/item?id=34874976