3 ms·
> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoin
by DeepSeaTortoise 26d ago
> If OpenAI is indeed using customer data to train their models to win a $1m prize
Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
- rickdeckard 26d agoAgree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results... [0] https://www.axios.com/2026/08/17/google-spirit-airlines-bankruptcy https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
- leonidasrup 26d ago> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
- rickdeckard 26d agoThe fantastic grey area that was engineered over the past decade is "profiling", so my guess is the answer will be "we didn't use your customer data for training, but we cannot rule out that it has been used to create profiles of your customers to train our model"
- marcosdumay 26d agoIf you still had that question, you can answer it now. But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.
- AlexCoventry 26d agoI'm not a lawyer, but I think the comparison to copyright law is invalid. Such a dispute would be governed by contract law.
- deleted 26d ago[deleted]
- chinathrow 26d ago> The chances of getting caught are 0 I'd say non-zero, as seen in the current state of affairs.
- bryanrasmussen 26d agosort of agree, but also sort of think an accusation with lots of people arguing is not exactly the same as being caught.
- barrkel 26d agoThis would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.
- lou1306 26d agoOk there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
- desterothx 26d agoBillions lost, while waiting for their trillion ipo. Im sure they would manage...
- theturtletalks 26d agoWhen these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them. But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
- noir_lord 26d agoAnd require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.
- DeepSeaTortoise 26d agoSure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves. You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training). Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise). In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with. So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data. The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it. And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with. So I'd say a whistleblower is pretty out of luck even becoming one.