5 ms·
You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything
by 20k 1mo ago
You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy
The idea that they'll steal from everyone except you is just wishful thinking
- WarmWash 1mo agoIt would be catastrohic for any of the big labs if it came out that they were training on what was sold as private. I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.
- r_lee 1mo agoexactly, I don't understand how HN doesn't understand this by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"
- 20k 1mo agoI mean, they've already violated the law in acquiring all their training data already, why would they be uncomfortable violating a contract to get more training data?
- vineyardmike 1mo agoBecause violating the law was the prerequisite to starting their business, without it they're worth $0 and have no models. They've already survived Training on customer input may help the models but it isn't "bet the farm" helpful. Now, they have a thriving business, so they shouldn't risk their business for incremental data that they can buy. At this point, the reputation of the business matters too. Fable explicitly didn't support zero-retention usage, and it saw significantly lower adoption vs other flagship models, and their past releases. Being caught abusing enterprise contracts is really hard to dig out of.
- r_lee 1mo agoand they're under no contract with those millions of websites and books, but it's a whole other thing when it's a paying customer, especially a large enterprise with a signed contract with DPAs etc.
- ajb 1mo agoIt would be catastrophic if they violated confidentiality blatantly, but that doesn't exclude learning of any description. For example, an ordinary human being can't fork a subagent for a particular client and wipe it afterwards. Humans can't stop themselves learning, so confidentiality can't ban all learning. Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes. So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time. Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it? Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.
- 20k 1mo agoI mean the models have literally trained on: 1. Child porn 2. Stolen music 3. Private github repos, before that was 'stopped' 4. Illegally pirated books Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?
- Henchman21 1mo agoThese folks need to feel repercussions so hard their souls flee to the afterlife leaving only their sad, dead husks behind.
- realusername 1mo agoIt's going to be the same as PRISM, people will be outraged and business will go back as usual. (And they totally won't do it again they swear, the contract says so)
- thephyber 1mo agoNo it wouldn't. It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts. When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes. Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file. And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.
- jppittma 1mo agothat was accidental? this would be straight up fraud.
- chillfox 1mo agoYou can structure it so that it becomes accidental. 1. Ensure security barriers are weak or honor based. 2. Put individual researchers under a lot of pressure. 3. If you get caught, blame the weak barriers, or the individual researcher. Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room with a sign on the door saying please don't enter.
- thephyber 1mo agoYeah, the steps follow exactly what happened at VW with DieselGate. The diesel emissions lies were found out because some enterprising person set up an emissions testing system and drove the car in real world scenarios with it to verify the claimed emissions. There's no reliable way to verify a foundation model has been trained on a particular piece of proprietary data. If an API key is ingested, hopefully the foundation model is wrapped in enough moderation that the raw API key oberserved during training is not recited verbatim in the output.
- tripledry 1mo agoI think it's arguable that it wouldn't be catastrophic. Doesn't have to be outright lying, it can be something like "opt out" being off by default, and them opting you in at the next update without you noticing. This is just a gut feeling and goes into the conspiratorial energy, but I'm pretty sure I can find countless examples of companies behaving like this in the past where it wasn't catastrophic (google meta amazon adobe ...)
- gitgud 1mo agoIf you have an enterprise contract with them, the legal protections for the consumer is much higher… is what I’ve been told anyway
- olejorgenb 1mo agoNot sure what you are saying exactly, but > the legal protections for the consumer is much higher You know you give away the right to file class action law suits against Anthropic when you accept their Terms Of Use, right? (at least the Americans ones)
- stingraycharles 1mo agoIs that the case with enterprise contracts as well? Can’t imagine any decent procurement / legal team accepting this.
- 20k 1mo agoIt isn't, in both cases you're protected by exactly the same legal system
- gitgud 1mo agoSame legal system, but an OpenAI enterprise contract is not the same as a ChatGPT pro subscription... The terms of service agreement would be quite different
- zamadatix 1mo agoIt's also the same planet but finding a commonality somewhere in the chain is not the same thing as finding a lack of difference in the matter.
- autoexec 1mo agoIf you have an enterprise contract with them you're still unlikely to ever know that you've been lied to unless some whistleblower at the AI company comes forward and even if you do somehow find out, it's too late. Once they have your data and have trained on it there's no taking it back. At most they'll pay out some tiny settlement that's a fraction of how much money they make in a week and they'll continue to profit from your data forever.
- solaire_oa 1mo agoPinky-promises, embarrassing. I'd point to tinfoil.sh. I'm not a shill for tinfoil, I haven't even used it or looked past its homepage really, but if we are able to legitimately secure privacy, then training data questions are moot (and the market for training data would probably shift/expose).
- lukewarm707 1mo agoagreed. the pricing is a nightmare. either $20 for codex and up to 00 millions of tokens or $0.50 cache in per million in tee.
- teeray 1mo agoThis is why these companies paying subscriptions shoveling everything into Claude thinking "oh, they're not training on our stuff" is hilarious to me. Of course they are. They probably are extra super-duper sure not to have Claude admit to that in any way, but there's no way they're giving up training on the sum total of both open and closed source code out there.
- vorticalbox 1mo agoGood example of this is figma. Claude was updated to work with it and now we have Claude design.
- someothherguyy 1mo agohttps://beckreedriden.com/the-black-box-problem-in-ai-trade-secret-litigation-how-do-you-prove-use/ https://beckreedriden.com/the-black-box-problem-in-ai-trade-...
- addag 1mo agoAnd even if they are not doing it right now, they probably keep the history available for future training, "just in case".
- protocolture 1mo ago>You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy Have been having a think about this statement. I agree in intent, some of them are probably breaking the agreement for training data. I dont think Microsoft is doing it, Enterprise Data Protection is the plank holding up their entire Copilot line. One whiff and everyone's gone. Copilot isnt actually good at anything except giving some illusion of protection, and preventing users from following a desire path to other LLMs without enterprise data protection. Its the core value proposition. But we only need to wait and see what the next 20 data breaches tell us to find out for sure. In detail however, I dont know if they could just claim it as fair use, after exclaiming that they specifically wont do that. I dont think "Fair Use" would be the issue so much as contract law. I know a EULA wouldnt hold up the other way (By reading this you agree not to steal my data and train an LLM with it) for data thats made freely available on the internet. But if you are purchasing the "No Training" contract they would be in breach if they trained with it. Possibly fair use would let them keep the data after paying whatever they owe in terms of contract breach.
- classified 1mo agoI really don't get how there could be anyone to whom this is not beyond obvious. And yet, software companies are donating their means of production to these crooks left and right, hastening the day where they really will become obsolete. I'm sure others do equally unwise things. If you're big enough, crime pays exceedingly well.
- j4k0bfr 1mo agoI think innocent-until-proven-guilty is the correct approach in general but... It's also immature to assume your data will stay private in the long term. These AI companies are immensely valuable targets. When they get breached, their datasets will inevitably penetrate public datasets. And why would any AI company refuse to train on 'public' data?
- PunchyHamster 1mo agoinnocent until proven guilty is definitely wrong approach vs tech giants who time and time again prove they will do anything unethical if only they can theoretize how to get away with it or the fine is low enough
- deadbabe 1mo agoI don’t care about what they do, I only care about what they say, that way, when compliance requires us to use LLMs that don’t train on inputs, I can point to that policy and continue on using the LLM. If they are secretly scraping input prompts, that no longer becomes my problem. Someday, a massive lawsuit comes down, and everyone gets to play the victim. You’re naive for thinking we don’t know how this really works.
- flakeoil 1mo agoWhat they say is not the same as what they wrote in the ToS you didn't read. If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from?
- rcxdude 1mo agoHave you read their ToS? It's not that hard. This feels like this nihilistic 'law is witchcraft and the big company lawyers always have you screwed in the fineprint' when it generally is pretty comprehensible with not that much work.
- deadbabe 1mo agoBecause I can get on with my life and do my job. Let the legal departments worry about that.
- PunchyHamster 1mo ago> If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from? The point is that if someone accuses you of leaking data you can sue them for not abiding by contract, not that the data is not used for training
- rcxdude 1mo agoThis only makes sense if your model is 'the labs are just breaking the rules all the time' as opposed to 'the labs have a legal theory of defense for the one aspect of the business which is legally questionable'. The question of whether doing something their terms of service explicitly say they will not do opens them up to civil liability is a much more certain one than anything they are doing around copyright. If they are doing this and anyone can prove it, they are going to be sued very hard, and it will be a pretty straightforward case. (To put it another way: if the copyright arguments against training on data scraped from the internet fail, the big AI companies don't really have a business, so they are going to proceed on the basis that they do until someone forces the issue otherwise. They don't need to train on data from their customers, and that is an argument that is almost certain to fail in court. If you're carrying 100kg of cocaine in your car, speeding is a really dumb idea, but the analogous action here would be firing a machine gun into the air from the driver's seat)
- knollimar 1mo agoWell maybe they have another suspicipus theory for defense of your data. Like taking it, laundering it through a summarizer, and training on that. Surely that's less questionable than taking copyrighted stuff and saying "don't output over 15 words, never verbatim" pretty please.
- rcxdude 1mo agoCopyright has a lot of leeway for arguments about fair use. 'We won't use your data for training' in a contract has far fewer.
- AlexandrB 1mo agoIt doesn't matter. Even if they don't train on your inputs now they will "boil the frog" and do so in the future. What the user wants is irrelevant in the tech industry. > If they are doing this and anyone can prove it, they are going to be sued very hard, and it will be a pretty straightforward case. This is a joke, right? The usual settlement in these cases amounts to a few days worth of revenue.