6 ms·
OpenAI has temporarily stopped selling the Plus plan
- Danjoe4 4y agoThe comments are so ignorant it's unreal. I think the OpenAI team is doing their best. I've heard it's actually hard to scale an app to 100+ million users
- plutonorm 4y agoNah. Super easy I heard /s
- contemplatter 4y agoGPT-4 is the greatest software developer in the world. I'm sure that it can scale itself easily.
- smoldesu 4y agocommit hash: deadbeef commit author: GPT-4 commit message: ":rocket_emoji: Set all Azure clusters to autoscale in order to resolve flooded message queue"
- emptysuitcase 4y ago[dead]
- true_religion 4y agoI think it’s because people associate OpenAi with Microsoft so they think that one of the largest cloud computing providers in the world ought to be able to scale to these numbers.
- rvz 4y agoImagine people saying 'Google is doing their best' whilst GCP goes down every week with startups like Anthropic and others getting investment from Google. It is also like accepting the hilarious availability times of GitHub and then saying 'GitHub is doing their best' whilst they have an outage every week. Lets just say that a service that claims to scale to more than 50 million or 100 million users going down each week is not acceptable or being an apologist for large companies unable to scale with the so-called best engineers working for them.
- shapefrog 4y agoIts Open so you can just grab a copy and spin up your own version ... /s
- j16sdiz 4y agofrom the reddit comments: > I bought it and it was removed off of my account. I paid money for a feature I never got. Anyone else get this? It look like a very bad experience. if i were him, i would issue a credit card charge back
- orbz 4y agoI’d at least give them an opportunity to fix and make me whole before jumping to charge backs.
- bloppe 4y agoGPT3 has 175 billion parameters and Altman said 4 would use way more compute than 3. Let's just look at GPT 3. Each forward pass requires 2N=350B flops per token. The computational overhead of the attention mechanism is negligible but the memory overhead of the attention for all users is not, but we can ignore that for now. Let's assume each query involves ~200 tokens of combined input and output. That's 70T flops per query. Let's say the cost of compute is 1 cent per 100B flops (probably a lowball). That's $7 per query for compute alone. And that's just GPT 3.
- Robotbeat 4y agoI mean, that might be the cost for a onesie-twosie in Azure, but if you're microsoft and you own the hardware, the cost may be one or more orders of magnitude less than that. (Of course, there's just a limit on the number of GPUs that exist and that Nvidia can pump out.)
- bloppe 4y agoMicrosoft's profit margin for "intelligent cloud" in their most recent earnings was less than 50%. Impressive, but not nearly enough to make GPT subscriptions make financial sense.
- Robotbeat 4y agoAnd a lot of that is marketing, keeping unused capacity in reserve, etc. Microsoft can probably afford a very low price without losing money if they price the compute at-cost. Certainly less than $7/request. 70 teraflop per request, and an $8000 GPU, with $2k of server overhead, $10k in power and cooling costs, $20k total for 312 teraflops of compute, leased for 2 years gives you an at-cost hardware price of $0.00007/request: https://www.google.com/search?q=%2420000*70/(312*3600*24*365*2) https://www.google.com/search?q=%2420000*70/(312*3600*24*365... So even assuming just a 1% utilization rate, you’re still talking less than 1 cent per request. Remember, inference runs at reduced precision. FP16 (or even just 8 bit) is WAY cheaper than fp64 bit. More than an order of magnitude cheaper. 9.7Tflops vs 312Tflops, factor of 32 improvement. With sparse matrix and 8-bit precision, you can effectively get over 100x the performance on an A100 than with FP64, plus the memory requirements are more than an order of magnitude less. Factor of 128 improvement in processing speed plus at least a factor of 8-16 or more improvement in tokens per memory, depending on sparsity. (Sparsity and reduced precision do reduce performance for the same number of operations and weights, but the overall effect on performance from reducing precision is beneficial probably at least to 4-bit quantization… You get diminishing or negative returns if you go less than that with current architectures, but in principle we can keep going.) https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent... I think your compute cost figure is both very old and probably assumes double precision whereas GPT-3 uses at most FP16 at inference and GPT-4 may go beyond that.
- ren_engineer 4y agoHave to assume they literally can't get enough GPUs to respond to demand. Wonder if $20/mo is even profitable for them?
- bloppe 4y agoThere's no way it's profitable. I'd wager this will materially affect Microsoft's next earnings report. Hard to tell how open ai itself is affected financially by the explosive growth but a rude awakening is coming for many users and shareholders alike
- unmole 4y ago> wager this will materially affect Microsoft's next earnings report How?
- bloppe 4y agoThey're footing the bill for this whole operation. They provide the compute and storage and they're also so heavily invested in OpenAI after their recent $10B round that they can't offload the costs. See my other comment on this post for a conservative cost estimate for each query.
- chaxor 4y agoIt's not necessarily a bad move to get them back in the game after many years of being useless. Deepmind is one of only a few thin threads that may hold Google together in the coming years, but it really doesn't look good for them
- Rastonbury 4y agoIt won't show up on their financials I reckon, they made equity investment into OpenAI - their P&L won't show on Microsoft's. If they have a sweetheart deal with MSFT, Azure revenues aren't going to change. Even if they built a new DC or bought a boatload of GPUs for OpenAI, it will get mixed into their massive Capex and they probably are using existing infrastructure
- McDev 4y agoFrom my experience with their Plus plan GPT 4 usually just times out no matter what the prompt, forcing me to revert to 3.5. I'm not sure why I haven't asked for a refund yet.
- mike_hearn 4y agoMaybe it depends on time of day? I've been using GPT-4 mostly in the evenings CET and it's been responsive and fast the whole time, I've never seen a timeout.
- Veen 4y agoI find the OpenAI API Playground is more reliable. It’s less convenient than ChatGPT Plus, but I’ve not had many timeouts. Plus, you can always use Bing Chat. It’s more locked down, but depending on what you’re doing, it’s pretty good and backed by GPT-4.
- int_19h 4y agoThe Playground doesn't have GPT-4 API yet. As for Bing, despite the claims that it is GPT-4, it's clearly inferior to what OpenAI is offering. It's probably an older iteration of the model, and I wonder if it might also be a scaled-down version to run it cheaper at scale.
- Veen 4y agoThe Playground certainly does have GPT-4. I've been using it there for a couple of weeks. Maybe you need to have applied to the beta and been given access? https://ibb.co/f2KJm4z https://ibb.co/f2KJm4z As for Bing? Who knows what's going on behind the scenes. I've read that the "Creative" mode uses GPT-4 while the others might use a faster model, but that may well be nonsense.
- int_19h 4y agoAh, indeed, that access is still rather limited; most can only have a peek a GPT-4 through ChatGPT Plus. As for Bing, I don't know what's behind the scenes, but it does noticeably worse on tasks in "Creative" mode than ChatGPT backend in GPT-4 mode, in my limited experiments. For example, try this: > A is 1m left of B, B is 1m above C, D is 1m right of C, E is 1m below D, and E is 1m right of F. Where is F located relative to C? GPT-4 can usually solve this correctly. Bing is usually wrong even when it tries to solve it step by step (and it often won't unless you prompt it).
- dzsekijo 4y agoBest gimmick ever. This did work out on me to get me subscribed ;)
- shapefrog 4y agoOO I can still subscribe ... I have only used it once or twice in the last 6 months, but if I dont subscribe now I might not be able to ever again ...