20 ms·
If this is a "wake up call" - then your legal team needs immediate education. First - there is this - https://openai.com/policies/how-your-data-is-used-to-impr
by ghshephard 26d ago
If this is a "wake up call" - then your legal team needs immediate education.
First - there is this - https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/ https://openai.com/policies/how-your-data-is-used-to-improve... (linked from the Navier Stokes writeup)
I don't know how much more clearly they can write:
> When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models.
One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality.
The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
- pms 26d agoChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated: > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more The "Learn more" link takes you to the link you've shared.
- pesacharia 26d agoAs I understand it, there is substantial question as to whether that actually stops them training on your data, it just perhaps changes what derivative processes are applied and used.
- ghshephard 25d agoAll I can say with certainty that the legal team at our typical Multi-Billion dollar Silicon Valley company had zero faith in any licensing arrangements with Anthropic or OpenAI, regardless of they $$$ involved, and that even getting to the point where Amazon Bedrock on Dedicated GPUs (we're already a big AWS customer - so definite cost advantages to dealing with them) - took 4-6 months of legal review before we could allow our engineers to start using Claude and OpenAI coding agents. Still can't use Fable because of their Data Retention requirements.
- tpkm 26d agoThe training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
- carljungslabtek 26d agoThis is not true. They claim not to train by default for business and enterprise agreements, but for plus and pro plans they enable it by default and you can allegedly turn it off (I don’t trust them very much though, I’m sure there is something in the T&C saying they can modify that deal any time)
- ghshephard 25d agoAbsolutely 100% not true. I have colleagues in 4 "FAANG adjacent" companies plus the one I work at - zero of them have any faith in Enterprise Agreements from either OpenAI or Anthropic. There's a reason why people spend more $$$ with Data Bricks, Palantir, AWS Bedrock etc.. and don't even consider using Anthropic or OpenAI APIs directly - it's because those guarantees provide very little in the way of data-discovery, audit requirements, or liquidated damages should it ever be discovered there was data leakage. At least with these other companies, while the LD is likewise not great (typically limited to the amount of money you paid them) - you at least have some data-governance guarantees around running on dedicated hardware - no multi-tenancy, no third-party access outside of the AWS operators who keep the HW running - but are very much not in the business of looking at your data. I think this is mostly a function of what's at risk - when company valuations get into the 10s of billions of dollars, the risk of IP leaking into what could be seen as competitive companies (OpenAI/Anthropic would be happy to take over the world - I don't sense that AWS or Azure, are as ruthless in stepping on their customers business, unless of course they are a SAAS provider) is just too significant a liability to take - particularly when you can de-risk.