4 ms·
I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? I don't mean this as
by lynndotpy 20d ago
I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs?
I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.
- Gud 20d agoNo, that is not "common knowledge". You are supposed to be able to disable that unwanted feature.
- ForHackernews 20d agoI have no inside information, but I always assume the tickboxes that "disable ____ data" from Google/Facebook/OpenAI just disconnects it from your own account, not hides it from the provider.
- Aurornis 20d agoThe checkbox has an actual statement associated with it about what it does. You don’t need to assume anything.
- chii 19d agoThere's no independent verification of what that checkbox actually does. The company can say anything, and you are unable to verify that they actually do it. The only verification you could do so far is GDPR-style data export, and also the adherence to GDPR regulations (and even those might get skirted if they aren't operating in europe).
- taneq 20d agoAren’t they usually phrased very specifically as “we collect this data and use it to show you relevant ads, you can opt out of us showing you relevant ads”?
- zdragnar 20d agoSome offer zero data retention policies, but there can be weasel words. For example, on the individual pro plan, you can turn off the setting that lets them train models on your data, but they still have a section in their terms that allows them to evaluate your anonymized data for statistical and "research" purposes. You have to actually get a signed contract along with an enterprise plan that spells out exactly what they're going to use, and what settings enable what retention. https://privacy.claude.com/en/articles/10023548-how-long-do-you-store-my-data https://privacy.claude.com/en/articles/10023548-how-long-do-... (see the additional info section)
- jmalicki 20d agoOr just use Azure, AWS, etc. for Claude/ChatGPT inference, where the AI labs never even get your data in their data centers at all. You pay more for it, but if you care that much, use it.
- FuckButtons 20d agoSeems naive to think that those providers - who have a financial interest in selling the data - would not also try to weasel out of the precise definition of ‘zero’ retention.
- ssivark 20d agoWhat about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
- hazard 20d agoI worked very briefly at Baseten, and I can say that it was a perpetual annoyance (from an engineering perspective) that customers would complain about issues with their models but we couldn't actually see the inputs/outputs. I don't know about the other providers, but at Baseten they literally weren't stored anywhere.
- lynndotpy 20d agoI don't have any much exposure to the attitudes people have around them, and I haven't worked with them. So I can't really say
- jmalicki 20d ago> using open weights models AWS and Azure give you the same thing for Claude and ChatGPT, no need to be stuck with open weights. They might sometimes store some of it for other purposes (I don't know the specifics), but it is emphatically not being fed back to OpenAI or Anthropic.
- ljlolel 19d agoA provider can genuinely avoid storing inputs, as the Baseten engineer below describes. That is still different from proving what code received the prompt or protecting plaintext while it runs; I built TrustedRouter to separate ZDR, attestation, and confidential routes: https://trustedrouter.com/blog/attestation-is-all-you-need?utm_source=hackernews&utm_medium=organic_reply&utm_campaign=openrouter_conversations_202609&utm_content=20260913_hackernews_inference_provider_zdr_terms https://trustedrouter.com/blog/attestation-is-all-you-need?u...
- ssivark 19d ago> to separate ZDR, attestation, and confidential routes Could you please clarify what that means? Given what I've been searching for, I might in principle be part of your intended customer profile, but I can't figure out whether you are merely doing routing (alternative to OpenRouter) or also inference (alternative to the names I've mentioned above). If it's merely routing, then how do you protect me from any potential misbehavior on the part of the inference provider? Just feedback for what you're building, so please take this in a positive spirit... I'm an AI researcher and not quite an infra guy, and I'm making recommendations on token APIs for several less knowledgeable around me (I've gotten a few people set up with Baseten recently), and I couldn't figure out whether/why I would be interested in TrustedRouter. You should communicate the story better :-) EDIT: Here's what I now understand after some digging; please correct if wrong. There are some M token providers (not the names I listed above?) who provide cryptographic guarantees about inference services. But somebody still needs to verify what they do on each request. For an individual running a single harness, that harness would be a logical place to perform this verification if possible. For an org with N users each running their own harness, TrustedRouter solves the N*M problem and becomes the single gateway for trusted inference -- provided one somehow trusts/verifies TrustedRouter.
- Aurornis 20d ago> I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? The services have toggles to allow prompts to be used in the training set. There is a conspiracy theory that the toggle is a false distraction and they’re actually keeping everything, and that none of the employees involved will ever whistleblow this fact. Outside of Internet comment sections, I think most people assume these US-based companies are doing what they say. For enterprise use there are services like AWS Bedrock which have strict isolation guarantees. There are some people who still believe those guarantees are a lie, but once someone has reached that point I don’t think they trust anything that isn’t running entirely within their house. People in that category are a very small minority, but a very vocal minority.
- lynndotpy 20d agoThe impression I have (from interacting with people IRL using OpenAI and Anthropics offerings, and how they feel about the risks involved) is just the opposite. But we probably just have different life experiences.
- Aurornis 20d agoI can name groups of people I interact with who lean both ways. It’s still a commonly held belief that “Facebook sells your data” and it’s cool to be cynical about everything tech in many social scenes. Conceding that a tech company might be honest about something will get you classified as a bootlicker depending on who you talk to so the only winning move is to be super cynical. Among actual professionals I work with in tech and legal, almost nobody holds a belief that these companies are blatantly lying to their customers (and zero of their employees are whistleblowing it, while said companies also have employees trying to whistleblow AI safety on Twitter daily)
- lynndotpy 19d agoYes, I am talking about working professionals who use LLMs. Before this thread, I would have considered it surprisingly and singularly naïve if someone told me they trusted OpenAI. I still believe the common and correct take is that these companies are largely training on customer data against their consent. I don't think they are "blatantly" lying either, just normal bog-standard lying that we've all come to accept. It's a profitable and competitive tech company. We have already seen this lying. The toggles are opt-out, not opt-in. When you sign up, you agree to binding arbitration, which is effective for preventing lawsuits in the US. The toggles are regularly turned back on without our consent on ChatGPT and Claude. OpenAI's "don't train on my content" setting isn't even in the ChatGPT interface. As far as I know, they haven't suffered even a tiny controversy in public opinion over any of this at all. There's nothing to whistleblow about when it's public knowledge. How many of the people who checked those boxes have cryptographic proof they did it? How many of those people have opted out of the arbitration clause? How many of those people would be able to claim damages? Would the amount of people who satisfy all three questions be large enough to make it worth _not_ training on user data?
- cgio 20d agoI would wager that’s more acceptable if said learning is not in competition with the user. If they didn’t actually produce results but created the model only, then that could be advantageous for users too. But the moment they absorb your work to sell it, or for marketing, it’s a different moral ground.
- bluecalm 19d agoYou have some secret sauce. The model trains on it. Your competitor is solving a similar problem. The model "advantageously" helps them. Your competitor is happy and continues to pay for the subscription. Sam and Dario just resold your code. For what it's worth LLMs still suck at reproducing my little secret algorithm/implementation while being able to solve way harder problems. I have a good guess why that's the case.