18 ms·
> I have no idea how OpenAI can make money on this. I did some quick calculation. We know the number of floating point operations per token for inference is ap
by vishal0123 4y ago
> I have no idea how OpenAI can make money on this.
I did some quick calculation. We know the number of floating point operations per token for inference is approximately twice the number of parameters(175B). Assuming they use 16 bit floating point, and have 50% of peak efficiency, A100 could do 300 trillion flop/s(peak 624[0]). 1 hour of A100 gives openAI $0.002/ktok * (300,000/175/2/1000)ktok/sec * 3600=$6.1 back. Public price per A100 is $2.25 for one year reservation.
[0]: https://www.nvidia.com/en-us/data-center/a100/ https://www.nvidia.com/en-us/data-center/a100/
[1]: https://azure.microsoft.com/en-in/pricing/details/machine-learning/ https://azure.microsoft.com/en-in/pricing/details/machine-le...
- lumost 4y agoIsn’t it 2.25 per hour per a100?
- TheMagicHorsey 4y agoYes, he means 2.25 per hour with a 1 yr reservation.
- p1esk 4y agoYou can get A100 on Lambda Labs cloud for $1.1/hr ($8.8/hr per 8xA100) without any reservation.
- gyrovagueGeist 4y agoIts a good baseline, but I very much doubt that openAI is paying anywhere near the public cost for their compute allocation.
- lumost 4y agoDirect purchasing isn’t too much cheaper. An H100 costs 35k new. OpenAI and MS are probably getting those for around 16k about 1.82 per hour.
- minimaxir 4y agoIt's speculated that ChatGPT uses 8x A100s, which flips the conclusion. Although the ChatGPT optimizations done to reduce costs could have also reduced the number of GPUs needed to run it.
- pelasaco 4y agoI checked the price of a A100, and its costs 15k? Is that right?
- alchemist1e9 4y agoAnd $2.25 per hour on 1 year reservation means 8,760 hours x 2.25 = $19,710 rent for the year. Not a bad yield for the provider at all, but makes sense given overheads and ROI expected.
- pelasaco 4y agoyes, specially that you don't have to deal with buying it, maintaining it, etc...
- sroussey 4y agoNot sure why people are so scared of this (in general). Yes, it’s a pain, but only an occasional pain. I’ve had servers locked up in a cage for years without seeing them. And the cost for bandwidth has plummeted over the last two decades. (Not at AWS, lol)
- Sebb767 4y agoThe problem isn't the good times, the problem is when something happens in the middle of the night, when a RAM stick goes bad or when you suddenly need triple the compute power. Usually, you get to feel the pain when you need it the least. I'm hosting a lot of stuff myself on my own hardware, so I do sympathize with this argument, but in a time>>money situation, going to the cloud makes a lot of sense.
- osigurdson 4y agoThis would be a really fun optimization challenge for sure!
- freeqaz 4y agoIt's also worth mentioning that, because Microsoft is an investor, they're likely getting these at cost or subsidized. OpenAI doesn't have to make money right away. They can lose a small bit of money per API request in exchange for market share (preventing others from disrupting them). As the cost of GPUs goes down, or they develop at ASIC or more efficient model, they can keep their pricing the same and then make money later. They also likely can make money other ways like by allowing fine-tuning of the model or charging to let people use the model with sensitive data.
- whatshisface 4y agoTheir new AI safety strategy is to slow the development of the technology by dumping, to lower the price too much to fund bootstrapped competitors.
- Silverback_VII 4y agoI highly doubt it. OpenAI, Google and Meta are not the only ones who can implement these systems. The race for AGI is one for power and power is survival.
- nr2x 4y agoLLM can do amazing things, but it’s a basically just an autocomplete system. It has the same potential to take over the world as your phones keyboard. It’s just a tool.
- darklycan51 4y agoThey want this, the interview from their CEO sorta confirmed that to me, he said some crap about wanting to release it slowly for "safety" (we all know this is a lie). But he can't get away with it with all the competition in other companies coming on top of China, Russia and others also adopting AI development
- npunt 4y agoYeah we're in an AI landgrab right now where at- or below-cost pricing is buying marketshare, lock-in, and underdevelopment of competitors. Smart move for them to pour money into it.
- madelyn-goodman 4y agoI really wonder if one way they are able to make money on it is by monetizing all the data that pours into these products by the second.
- bboygravity 4y agoThe could probably live off of the NSA sponsoring alone.
- ddmma 4y agoSpot on
- drexlspivey 4y agothe only one making money on this is NVIDIA
- bigfudge 4y agoSelling shovels in the goldrush…
- ilaksh 4y agoThey also mention in the new API docs that they are no longer keeping data submitted to ChatGPT. Or at least not to the ChatGPT API.
- _just7_ 4y agoWould probably pile up to an inhuman amount of data storage. Imagine having to pay for storing the equivalent of 1000 tokens of text within that budget of only 0.0002 dollars
- Tepix 4y agoThat's one zero too many. Storage cost of 1000 tokens (6000 bytes) on a single HDD is $0.000000096 assuming $16/TB
- kkielhofner 4y agoBut those A100s only come by eight and it’s speculated the model requires eight (VRAM). For a three year reservation that comes to over $96k/yr - to support one concurrent request.
- ALittleLight 4y agoWhat do you mean one concurrent request? Can't you have a huge batch size to basically support a huge number of concurrent requests? e.g. Endpoint feeds a queue, queue fills a batch, batched results generate replies. You are simultaneously fulfilling many requests.
- kkielhofner 4y agoHopefully they’re doing plenty of batching - you don’t even need to roll your own as you’re describing. Inference servers like Triton will dynamically batch requests with SLA params for max response time (for example). That said I don’t think anyone anyone outside of OpenAI knows what’s going on operationally. Same goes for VRAM usage, potential batch sizes, etc. This is all wild speculation. Same goes for whatever terms OpenAI is getting out of MS/Azure. What isn’t wild speculation is that even with three year reserve pricing last gen A100x8 (H100 is shipping) will set you back $100k/yr - plus all of the usual cloud bandwidth, etc fees that would likely increase that by at least 10-20%. We’re talking about their pricing and costs here. This gives a general idea what anyone trying to self host this would be up against - even if they could get the model.
- vishal0123 4y ago> will set you back $100k/yr This is 6 month of salary of one average developer's salary there. And BTW they are likely doing inference on 100s or 1000s of GPUs, not just 8.
- kkielhofner 4y agoYes and a devops engineer to manage an even moderately complex cloud deployment is an average of an extra $150k/yr. I don't know where this "cloud labor skill, knowledge, experience, and time is free" thinking comes from. 8, 80k, or 800k GPUs depending on requirements and load - the point remains the same.
- Dave_Rosenthal 4y agoNote that they also charge equally for input and output tokens but, as far as I understand, processing inputs tokens is much computationally cheaper, which drops their price further.
- dharma1 4y agoReckon they will (if not already) use 4bit or 8bit precision and may not need 175b params
- lee101 4y ago[dead]
- cubefox 4y ago"We know the number of floating point operations per token for inference is approximately twice the number of parameters" Does someone have a source for this? (By the way, it is unknown how many parameters GPT-3.5 has, the foundation model which powers finetuned models like ChatGPT and text-davinci-003. GPT-3 had 175 billion parameters, but per the Hoffmann et al Chinchilla paper it wasn't trained compute efficiently, i.e. it had too many parameters relative to its amount of training data. It seems likely that GPT-3.5 was trained on more data with fewer parameters, similar to Chinchilla. GPT-3: 175B parameters, 300B tokens; Chinchilla: 70B parameters, 1.4T tokens.)
- vishal0123 4y agohttps://arxiv.org/pdf/2001.08361.pdf https://arxiv.org/pdf/2001.08361.pdf. See the C_forward formula approxiamtion.
- cubefox 4y agoThank you. Though it isn't quite clear to me whether the additive part is negligible?
- sytelus 4y agoFrom eq 2.2, additive part is usually in few 10s of millions. So, for N > 1B, approximation should be good but it doesn't work. For example, GPT3 inference flops is actually 3.4E+18 so the ratio is 19,000 not 2.
- vishal0123 4y agoFrom the paper > For contexts and models with d_model > n_ctx/12, the context-dependent computational cost per token is a relatively small fraction of the total compute. For GPT3, n_ctx is 4096 and d_model is 12228 >> 4096/12.
- smy20011 4y agoThe 600t performance is with sparsity in the spec. I think the price is nearly break even if sparsity is not used in the model.
- cavisne 4y agoDoes openai actually specify the size of the model? InstructGPT 2B outperformed gpt 3 175B, and chatgpt has a huge corpus of distilled prompt -> response data now. I’m assuming most of these requests are being served from a much smaller model to justify the price. OpenAI is fundamentally about training larger models, I doubt they want to be in the business of selling A100 capacity at cost when it could be used for training