5 ms·
"There was one surprise when I revisited costs: OpenAI charges an unusually low $0.0001 / 1M tokens for batch inference on their latest embedding model. Even co
by arkmm 1y ago
"There was one surprise when I revisited costs: OpenAI charges an unusually low $0.0001 / 1M tokens for batch inference on their latest embedding model. Even conservatively assuming I had 1 billion crawled pages, each with 1K tokens (abnormally long), it would only cost $100 to generate embeddings for all of them. By comparison, running my own inference, even with cheap Runpod spot GPUs, would cost on the order of 100× more expensive, to say nothing of other APIs."
I wonder if OpenAI uses this as a honeypot to get domain-specific source data into its training corpus that it might otherwise not have access to.
- deleted 1y ago[deleted]
- cedws 1y agoI don’t think OpenAI train on data processed via the API, unless there’s an exception specifically for this.
- dannyw 1y agoCan you truly trust them though?
- cedws 1y agoYes, it would be disastrous for OpenAI if it got out they are training on B2B data despite saying they don’t.
- mattigames 1y agoYeah, so many companies have been completely ruined after similar PR disasters /s
- j33zusjuice 1y agoTheir terms of service say they won’t use the data for training, so it wouldn’t just be a PR disaster; it’d be a breach of contract. They’d be sued into oblivion.
- reasonableklout 1y agoHave they said they don't? (actually curious)
- gkbrk 1y agoYes, they have. [1] > Your data is your data. As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us). [1]: https://platform.openai.com/docs/guides/your-data https://platform.openai.com/docs/guides/your-data
- dweinus 1y agoWe're both talking about the company whose entire business model is built on top of large scale copyright infringement, right?
- dymk 1y agoNot the same when the people you infringe on can sue you into the dirt
- johnthescott 1y agoi am too lazy to ask openai.
- trhway 1y agoi'd not rule out some approach like instead of training directly on the data, may be they would train on a very high dimensional embedding of such a data (or some other similarly "anonymized", yet still very semantically rich representation of the data)
- dpoloncsak 1y agoMaybe I misunderstand, but I'm pretty sure they offer an option for cheaper API costs (or maybe its credits?) if you allow them to train on your API requests. To your point, pretty sure it's off by default, though Edit: From https://platform.openai.com/settings/organization/data-controls/sharing https://platform.openai.com/settings/organization/data-contr... Share inputs and outputs with OpenAI "Turn on sharing with OpenAI for inputs and outputs from your organization to help us develop and improve our services, including for improving and training our models. Only traffic sent after turning this setting on will be shared. You can change your settings at any time to disable sharing inputs and outputs." And I am 'enrolled for complimentary daily tokens.'
- anothernewdude 1y agoIt'd be a way to put crap or poisoned data into their training data if that is the case. I wouldn't.
- magicalhippo 1y ago> OpenAI charges an unusually low $0.0001 / 1M tokens for batch inference on their latest embedding model. Is this the drug dealer scheme? Get you hooked later jack up prices? After all, the alternative would be regenerating all your embeddings no?