4 ms·
But the cost is _definitely_ falling. For a recent example, see DeepSeek V3[1]. It's a model that's competitive with GPT-4, Claude Sonnet. But cost ~$6 Million
by silveraxe93 2y ago
But the cost is _definitely_ falling. For a recent example, see DeepSeek V3[1]. It's a model that's competitive with GPT-4, Claude Sonnet. But cost ~$6 Million to train.
This is ridiculously cheaper than what we had before. Inference is basically getting an 10x cheaper per year!
We're spending more because bigger models are worth the investment. But the "price per unit of [intelligence/quality]" is getting lower and _fast_.
Saying that models are getting more expensive is confusing the absolute value spent with the value for money.
- [1] https://github.com/deepseek-ai/DeepSeek-V3/tree/main https://github.com/deepseek-ai/DeepSeek-V3/tree/main
- ADeerAppeared 2y ago> Inference is basically getting an 10x cheaper per year! You're gonna need some good citations for that. There's a big difference between companies saying "The inference costs on our service are down" and the inference costs on the model are down. The former is oft cheated by simplyifying and dumbing down the models used in the service after the initial hype and benchmarks. > But the "price per unit of [intelligence/quality]" is getting lower and _fast_. Absolutely not a general trend across models. At best, older models are getting cheaper to run. Newer models are not cheaper "per unit of intelligence". OpenAI's fany new reasoning models are orders of magnitude more expensive to run whilst being ~linear improvements in real world capabilities.
- silveraxe93 2y agoSee situational-awareness[1], see the "algorithmic efficiencies" section. He shows many examples of how models are getting cheaper. With many citations. Costs are not just down on a specific service. Even though I don't see the problem in that, as long as you get the promised level of performance, without being subsidised. See the deepseek model I linked above. It's an open model and you can run it yourself. > At best, older models are getting cheaper to run. What's your definition of old here? If you compare the literal bleeding edge model (o3) to 2 years ago best model (GPT-4)? Not only is this a ridiculously misleading comparison, it's not even valid! o3 is a reasoning model. It can spend money at test time to improve results. Previous models don't even have this capability. You can't look at one example of where they just threw a lot of money and say this is the cost. The cost is unbounded! If they want, they can just not let the model think for ages and have basically "0-thinking" outputs. This is what you use to compare models. If you compare _todays_ cost for training and inference of a model as good as GPT-4 when it was released, this cost has massively gone down on both counts. [1] - https://situational-awareness.ai/from-gpt-4-to-agi/#The_trends_in_deep_learning https://situational-awareness.ai/from-gpt-4-to-agi/#The_tren...
- mvdtnz 2y ago> We're spending more because bigger models are worth the investment Are they? Where's the value? What are they being used for actually out there in the real world? Not the shitty apps that simonw bleats about day in day out, not the lame website bots that repeat your FAQ back at me - actual real valuable (to the tune of the billions being invested in them) use cases?
- silveraxe93 2y agoChatGPT is one of the fastest growing apps ever. Saying that's there's no products is willful blindness by this point. This is hackernews. I'd expect users to have a basic understanding of VC investment. The expected value of next-gen models times the probability to create them is higher than the billions than they are throwing at it.
- hatefulmoron 2y ago> ChatGPT is one of the fastest growing apps ever. Saying that's there's no products is willful blindness by this point. That's fair, but I think you're being a little uncharitable to the point being made. I would postulate that most ChatGPT users are not using it in a productive capacity, they're using it as a sort of "google that's better at understanding my queries." Obviously that serves a great niche for lots of people, but I don't think it's what mvdtnz had in mind.
- KaiserPro 2y agoI'm not convinced about that 10 cheaper a year. Larger models need more memory. I'm willing to bet that most of the tier 1 providers rely on multi-GPU models to serve traffic. None of that is cheap, 8x GPU nodes that serve less than 20 queries a second are exceedingly expensive to run.
- silveraxe93 2y agoLarger models are more expensive to run (ceteris paribus). But we're seeing we can squeeze more performance from smaller models. You need to compare like-for-like. You can't say that the cost of building a 5-story apartment is increasing by pointing at the burj khalifa.
- menaerus 2y agoNow remind us what HW did we need to run local inference of llama2-69B (July, 2023)? And then contrast it to the HW we need to run llama3.1-70B (July, 2024)? In particular, which optimizations and in what way did they dramatically cut down the cost of the inference? I seriously don't get this argument and I see it being repeated all over and over again. Although model possibilities are increasing, no doubt in that, HW costs for inference remained the same and they're mostly driven by the amount of (V)RAM you need.