4 ms·
Yes, openAI is dumping the market with chat-gpt 3.5. Vulture capital behaviour at its finest, and I'm sure government regulations will definitely catch on to t
by mrybczyn 3y ago
Yes, openAI is dumping the market with chat-gpt 3.5. Vulture capital behaviour at its finest, and I'm sure government regulations will definitely catch on to this in 20 or 30 years...
It's cheaper than the ELECTRICITY cost of running a llama-70 on your own M1.Max (very energy efficient chip) assuming free hardware.
I guess they are also getting a pretty good cache hit rate - there are only so many questions people ask at scale. But still, it's dumping.
- PUSH_AX 3y agoYou think they are caching? Even though one of the parameters is temperature? Can of worms, and should be reflected in the pricing if true, don't get me started if they are charging per token for cached responses. I just don't see it.
- why_only_15 3y agoYou can keep around the KV cache from previous generations which lowers the cost of prompts significantly.
- read_if_gay_ 3y agoturbo is likely nowhere near 70b.
- sacred_numbers 3y agoBased on my research, GPT-3.5 is likely significantly smaller than 70B parameters, so it would make sense that it's cheaper to run. My guess is that OpenAI significantly overtrained GPT-3.5 to get as small a model as possible to optimize for inference. Also, Nvidia chips are way more efficient at inference than M1 Max. OpenAI also has the advantage of batching API calls which leads to better hardware utilization. I don't have definitive proof that they're not dumping, but economies of scale and optimization seem like better explanations to me.
- hutzlibu 3y agoI also do not have proof of anything here, but can't it be both? They have lots of money now and the market lead. They want to keep the lead and some extra electricity and hardware costs are surely worth it for them, if it keeps the competition from getting traction.
- csjh 3y agoWhat makes you think 3.5 is significantly smaller than 70B?
- haxton 3y agogpt3.5 turbo is (mostly likely) Curie which is (most likely) 6.7b params. So, yeah, makes perfect sense that it can't compete with a 70b model on cost.
- ronyfadel 3y agoIt still does a much better job at translation than llama 2 70b even, at 6.7b params
- two_in_one 3y agoIf it's MOE that may explain why it's faster and better...
- yumraj 3y agoMOE?
- sarthaksrinivas 3y agoMixture of Experts Model - https://en.wikipedia.org/wiki/Mixture_of_experts https://en.wikipedia.org/wiki/Mixture_of_experts
- csjh 3y agoIs there a source on that? I've never seen anyone think it's below even 70B
- why_only_15 3y agogpt3.5 turbo is a new model, not Curie. As others have stated, it probably uses Mixture of Experts which lowers inference cost.
- jiggawatts 3y agoI thought it was fairly well established that GPT 3.5 has something like 130B parameters and that GPT 4 is on the order of 600-1,000