6 ms·
This is basically the end of OpenAI hardware. This is by far worst than the Waymo vs. Uber lawsuit which killed the Uber self driving project. Also if you are
by impulser_ 3mo ago
This is basically the end of OpenAI hardware. This is by far worst than the Waymo vs. Uber lawsuit which killed the Uber self driving project.
Also if you are a business using OpenAI models, I would highly suggest you do not because they are most likely looking at your code and IP.
- glenpierce 3mo agoGood point. If these are the ethical standards they go by, who’s to say they’re abiding by any standards to keep my data private.
- tiahura 3mo agoMy guess it winds up like the old FAT joke about Android.
- nullbio 3mo agoObviously they are looking at your IP and code. Anthropic trains on your data regardless of you opting out, I know that one for certain. There's no coincidence they "keep your data temporarily despite opting out" - because they wash it in legal loopholes. There is no opt out. These companies WILL steal your business. Only a matter of time before they are sued as well.
- Chu4eeno 3mo agoWeren't they required by a court to keep everything, despite the privacy policy etc.? Or am I mixing up my companies in constant law battles.
- SpicyLemonZest 3mo agoThey were required to keep all non-European user logs for a temporary period between April and September 2025, because the media companies suing them think these logs may be the only evidence in existence that could prove or disprove their alleged misconduct.
- azinman2 3mo ago> Anthropic trains on your data regardless of you opting out, I know that one for certain. How do you know that?
- MagicMoonlight 3mo ago[dead]
- general1465 3mo agoThey systematically violated copyright when they grabbed whole internet to train their models. Do you really believe that they will stop stealing because they signed some funny ToS? Especially when every bit of data they have and competition does not have is making their model better.
- cromka 3mo agoPeople downvote you like you're being paranoid, but we're literally discussing this in a thread that shows how little respect those companies have for any sort of trade secrets.
- nullbio 3mo ago[dead]
- user43928 3mo agoHow does the allegation make any sense? AI labs can hardly just throw random confidential data into the training and then hope it does not leak into the output of their model in an obvious way. If that would be found it would destroy their main source of revenue, it could became a major national security or healthcare enforcement matter, and result in criminal investigations.
- Barbing 3mo agoSome of the smartest people on the planet all in the same room, data at their fingertips… they randomly add it to the training set? Labs at least must study prompts in an airgapped fashion. From there, consider how they could generate synthetic data to train another model. After, require trusted staff to do multiple levels of independent granular reviews of all fruits of the highest-value stolen inputs. (Or for model training data only, data never has to leave the airgap.) Definitely risky, anyway. Surely some AI user has sent data, in confidential mode, with a unique shape they expect to be able to recognize if a later model recreated a facsimile even with heavy substitutions… but labs could bring risk of getting caught (over next few years) down quite low with extraordinarily ultraparanoid strategy. (But hopefully everybody is just behaving!)
- hhh 3mo agoHow is it obvious? We have strong legal agreements that state otherwise, do you think they are just lying and risking thousands of lawsuits? I think it's more likely that there are 3/4 of a billion users that don't have these agreements and just pay for ChatGPT Plus and don't opt-out of anything, and are feeding the scaling machine every day.
- nullbio 3mo ago> do you think they are just lying Yes. They're constantly lying, and constantly getting caught for it. They have a reputation for it. Why do you think this would be any different? Their standard opt-out agreement frames it as if they won't train on your data, but they do anyway, due to legal loopholes. They essentially clean-room everyone who opts-out, so while it's "technically" not training on "your" data, to the model it makes no difference. Your alpha and IP is not safe. Paying customers are now more easily able to clone your business as well, not just Anthropic themselves. The only reason this hasn't leaked yet is fear. Anthropic is a very litigious and dangerous company. Only a matter of time though, someone there will grow a spine and speak up.
- aesthesia 3mo agoCan you elaborate on the loopholes here?
- nsagent 3mo agoI'm unwilling to speculate whether or not OpenAI is breaking their agreements (I honestly have no clue), but as an NLP researcher I'm certain they could launder data by having an LLM rewrite it and subsequently train on the rewritten data. Papers like "Curated Synthetic Data Doesn't Have to Collapse" [1] and "How to Synthesize Text Data without Model Collapse?" [2] demonstrate it's possible to do this. Since OpenAI's Privacy Policy [3] explicitly allows for the use of deidentified data, it's possible they consider rewrites (maybe paired with a model used to identify explicit PII) to be deidentified. Whether OpenAI's legal team thinks rewriting in this way technically means they aren't training on your data isn't something I'm able to comment on. Here's the relevant Privacy Policy statement: We also aggregate or de-identify Personal Data so that it no longer identifies you and use this information for the purposes described above, such as to analyze the way our Services are being used, to improve and add features to them, and to conduct research. We will maintain and use de-identified information in de-identified form and not attempt to reidentify the information, unless required by law. Please note all the hedging words I used (maybe, possibly, etc). I honestly have no clue if they are doing this. I'm merely elaborating on a possible loophole like you asked. [1]: https://arxiv.org/abs/2605.07724 https://arxiv.org/abs/2605.07724 [2]: https://arxiv.org/abs/2412.14689 https://arxiv.org/abs/2412.14689 [3]: https://openai.com/policies/privacy-policy/ https://openai.com/policies/privacy-policy/
- everfrustrated 3mo ago>Anthropic trains on your data This is why companies are wanting the AWS hosted models because they trust AWS running of the same models more than the vendors themselves.
- frenchtoast8 3mo agoOnly for Opus etc. For “Fable 5, Mythos 5, and future models on Bedrock with similar or higher capability levels,” your usage data is retained for 30 days and shared with Anthropic which you can’t opt out of.
- angulardragon03 3mo agoThis is why we don’t use Mythos/Fable-series models at $WORK (hosted on Vertex).
- nullbio 3mo agoPlaying with fire.
- sk4rekr0w 3mo agoThis is way overstated. OpenAI will certainly launch devices. It is to be seen how competitive they are and how much product market fit they achieve. OpenAI also has better data retention policies relative to Anthropic on SOTA models.
- alpineman 3mo agoDo you really believe those policies after stories like this?
- sk4rekr0w 3mo agoThe word "stories" in your sentence is very instructive
- samtheprogram 3mo agoIt's a "news story". Do you think Apple would be suing OpenAI if it was just a "story"?
- smith7018 3mo agoIt's also not just a story; it's a lawsuit with evidence.
- overgard 3mo agoUse local hardware if you can people! Chatbots are a luxury anyway and the local ones are catching up. You don't need the fanciest bells and whistles. Don't believe the narrative that you're "falling behind" if you're not using the latest model 20 hours a day.