5 ms·
Obviously they are looking at your IP and code. Anthropic trains on your data regardless of you opting out, I know that one for certain. There's no coincidence
by nullbio 3mo ago
Obviously they are looking at your IP and code. Anthropic trains on your data regardless of you opting out, I know that one for certain. There's no coincidence they "keep your data temporarily despite opting out" - because they wash it in legal loopholes. There is no opt out. These companies WILL steal your business. Only a matter of time before they are sued as well.
- Chu4eeno 3mo agoWeren't they required by a court to keep everything, despite the privacy policy etc.? Or am I mixing up my companies in constant law battles.
- SpicyLemonZest 3mo agoThey were required to keep all non-European user logs for a temporary period between April and September 2025, because the media companies suing them think these logs may be the only evidence in existence that could prove or disprove their alleged misconduct.
- azinman2 3mo ago> Anthropic trains on your data regardless of you opting out, I know that one for certain. How do you know that?
- MagicMoonlight 3mo ago[dead]
- general1465 3mo agoThey systematically violated copyright when they grabbed whole internet to train their models. Do you really believe that they will stop stealing because they signed some funny ToS? Especially when every bit of data they have and competition does not have is making their model better.
- cromka 3mo agoPeople downvote you like you're being paranoid, but we're literally discussing this in a thread that shows how little respect those companies have for any sort of trade secrets.
- nullbio 3mo ago[dead]
- user43928 3mo agoHow does the allegation make any sense? AI labs can hardly just throw random confidential data into the training and then hope it does not leak into the output of their model in an obvious way. If that would be found it would destroy their main source of revenue, it could became a major national security or healthcare enforcement matter, and result in criminal investigations.
- Barbing 3mo agoSome of the smartest people on the planet all in the same room, data at their fingertips… they randomly add it to the training set? Labs at least must study prompts in an airgapped fashion. From there, consider how they could generate synthetic data to train another model. After, require trusted staff to do multiple levels of independent granular reviews of all fruits of the highest-value stolen inputs. (Or for model training data only, data never has to leave the airgap.) Definitely risky, anyway. Surely some AI user has sent data, in confidential mode, with a unique shape they expect to be able to recognize if a later model recreated a facsimile even with heavy substitutions… but labs could bring risk of getting caught (over next few years) down quite low with extraordinarily ultraparanoid strategy. (But hopefully everybody is just behaving!)
- user43928 3mo agoThat is an interesting thought. They could run some sort of analysis to find high value input, such as proprietary technology, algorithms, or strategy. Then they could group them together for one specific topic, and produce a report that analyzes if the information is plausible. If so, they can have it send to staff for review, who could then create a test set that rewards the model for going into the direction of the proprietary solutions known to work. I'm no expert, but at least something like that sounds plausible to me. I still very much doubt they are doing this.
- danshipt 3mo ago[flagged]
- latexr 3mo agoYour parent comment isn’t saying they doubt the assertion, they’re asking for details. If you say you know for certain, it makes sense to ask how. It makes a big difference if the answer is “I used to work there”, or “I implemented those systems myself”, or “I heard my cousin’s second ex-wife say she heard it from her hairdresser”, or “aliens visited me in my dreams and told me”. I don’t doubt these companies are lying through their teeth. We have plenty of proof of several cases where they did, to the point believing they are liars is a sensible default, but still I could not say I know for certain of every instance of their lies. Knowing how empowers you to do something about it and convince others.
- nullbio 3mo agoNo one is going to admit to having an inside source on HackerNews. Read between the lines.
- latexr 3mo ago> No one is going to admit to having an inside source on HackerNews. Not only is that not true (people make throwaway accounts specifically to share insider info), no one has said this was insider information, there are plenty of other ways to know these details. > Read between the lines. That means nothing. There’s no information given, there’s nothing to read between.
- vorticalbox 3mo agoWell we don’t know but just look at figma, claude for it happened and then Claude design come out. They knew exactly how developers worked from using figma as training data.
- llelouch 3mo agoProfessionals still use figmas. Claude design jsut expanded the space. Non-design specialized use Claude design now.
- preg_match 3mo agoSure but the intention was to compete with Figma, and they are to a degree. The success of their tomfoolery is really a separate issue from if they are engaging in said tomfoolery, which I think they are. I also think anyone expecting honestly from this is naive. Bad people doing bad things are not stupid, in fact they're usually pretty smart. Nobody is gonna come out and say what's going on, because that's purely self-destructive. We will get leaks and lawsuits slowly, much like OpenAI.
- hhh 3mo agoHow is it obvious? We have strong legal agreements that state otherwise, do you think they are just lying and risking thousands of lawsuits? I think it's more likely that there are 3/4 of a billion users that don't have these agreements and just pay for ChatGPT Plus and don't opt-out of anything, and are feeding the scaling machine every day.
- nullbio 3mo ago> do you think they are just lying Yes. They're constantly lying, and constantly getting caught for it. They have a reputation for it. Why do you think this would be any different? Their standard opt-out agreement frames it as if they won't train on your data, but they do anyway, due to legal loopholes. They essentially clean-room everyone who opts-out, so while it's "technically" not training on "your" data, to the model it makes no difference. Your alpha and IP is not safe. Paying customers are now more easily able to clone your business as well, not just Anthropic themselves. The only reason this hasn't leaked yet is fear. Anthropic is a very litigious and dangerous company. Only a matter of time though, someone there will grow a spine and speak up.
- aesthesia 3mo agoCan you elaborate on the loopholes here?
- nsagent 3mo agoI'm unwilling to speculate whether or not OpenAI is breaking their agreements (I honestly have no clue), but as an NLP researcher I'm certain they could launder data by having an LLM rewrite it and subsequently train on the rewritten data. Papers like "Curated Synthetic Data Doesn't Have to Collapse" [1] and "How to Synthesize Text Data without Model Collapse?" [2] demonstrate it's possible to do this. Since OpenAI's Privacy Policy [3] explicitly allows for the use of deidentified data, it's possible they consider rewrites (maybe paired with a model used to identify explicit PII) to be deidentified. Whether OpenAI's legal team thinks rewriting in this way technically means they aren't training on your data isn't something I'm able to comment on. Here's the relevant Privacy Policy statement: We also aggregate or de-identify Personal Data so that it no longer identifies you and use this information for the purposes described above, such as to analyze the way our Services are being used, to improve and add features to them, and to conduct research. We will maintain and use de-identified information in de-identified form and not attempt to reidentify the information, unless required by law. Please note all the hedging words I used (maybe, possibly, etc). I honestly have no clue if they are doing this. I'm merely elaborating on a possible loophole like you asked. [1]: https://arxiv.org/abs/2605.07724 https://arxiv.org/abs/2605.07724 [2]: https://arxiv.org/abs/2412.14689 https://arxiv.org/abs/2412.14689 [3]: https://openai.com/policies/privacy-policy/ https://openai.com/policies/privacy-policy/
- everfrustrated 3mo ago>Anthropic trains on your data This is why companies are wanting the AWS hosted models because they trust AWS running of the same models more than the vendors themselves.
- frenchtoast8 3mo agoOnly for Opus etc. For “Fable 5, Mythos 5, and future models on Bedrock with similar or higher capability levels,” your usage data is retained for 30 days and shared with Anthropic which you can’t opt out of.
- angulardragon03 3mo agoThis is why we don’t use Mythos/Fable-series models at $WORK (hosted on Vertex).
- nullbio 3mo agoPlaying with fire.