5 ms·
I work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable sce
by esalman 25d ago
I work in catastrophe risk modeling and it's a multi billion dollar industry.
We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models.
If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
- DeepSeaTortoise 25d ago> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
- rickdeckard 25d agoAgree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results... [0] https://www.axios.com/2026/08/17/google-spirit-airlines-bankruptcy https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
- leonidasrup 25d ago> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
- rickdeckard 25d agoThe fantastic grey area that was engineered over the past decade is "profiling", so my guess is the answer will be "we didn't use your customer data for training, but we cannot rule out that it has been used to create profiles of your customers to train our model"
- marcosdumay 25d agoIf you still had that question, you can answer it now. But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.
- AlexCoventry 25d agoI'm not a lawyer, but I think the comparison to copyright law is invalid. Such a dispute would be governed by contract law.
- deleted 25d ago[deleted]
- chinathrow 25d ago> The chances of getting caught are 0 I'd say non-zero, as seen in the current state of affairs.
- bryanrasmussen 25d agosort of agree, but also sort of think an accusation with lots of people arguing is not exactly the same as being caught.
- barrkel 25d agoThis would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.
- lou1306 25d agoOk there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
- desterothx 25d agoBillions lost, while waiting for their trillion ipo. Im sure they would manage...
- theturtletalks 25d agoWhen these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them. But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
- noir_lord 25d agoAnd require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.
- DeepSeaTortoise 25d agoSure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves. You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training). Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise). In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with. So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data. The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it. And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with. So I'd say a whistleblower is pretty out of luck even becoming one.
- steve1977 25d agoPretty much everything or at least a lot of what you use as an OpenAI (or Anthropic or whatever) customer was once someone elses product that just got appropriated by OpenAI. Why should your business be any different?
- motbus3 25d ago> We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. I feel this is exactly what will happen as they cause all sites to go closed source to protect their intellectual property and the AI companies offer only biased information. They are replace the business on internet model by bankrupting everyone with their own tools. This is predatory pricing under most antitrust laws (imho, not a lawyer) and it is very easy to do when you dont need to pay for the raw material.
- KoolKat23 25d agoThis is the business case already. And has been the case with tech companies for a long time. Your phones built in photo manager replaced a lot of what Photoshop does.
- noir_lord 25d ago> If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses. I mean how could you expect them to not given they've trained the existing models on effectively the sum total of all human knowledge available on the internet without regard to copyright/ownership of that material. It's a little trite but this absolutely runs into the "Frog and the Scorpion", it is simply in their nature.
- KoolKat23 25d agoI mean this is a basic question, is your data used to improve models, did you opt in or not. There's nothing crazy here and it's not identifiable. You'll just conveniently find the next model iteration knows how to do it. Any enterprise worth their salt already considers this stuff.