5 ms·
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' po
by torginus 14d ago
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
- ronsor 14d agoAlmost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data.
- applfanboysbgon 14d agoZDR is based on the exact same pinky-promise as training opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data.
- pennomi 14d agoWhere nothing is retained at all, allegedly.
- steveBK123 14d agoThey already trained on pirated content, what makes you think they are going to honor ZDR?
- GrinningFool 14d agoContractual obligations carry teeth. Scraping the internet is relatively risk-free.
- steveBK123 14d agoGood luck proving your data was laundered and included in a training run
- keeda 14d agoLucky for us Apple is already alleging something to this effect in their trade secret lawsuit, so you know they'll make sure discovery turns this up if it exists.
- ForHackernews 14d agoI mean... https://www.wsj.com/tech/ai/jury-sides-with-openai-sam-altman-in-case-brought-by-elon-musk-933240ff https://www.wsj.com/tech/ai/jury-sides-with-openai-sam-altma...
- applfanboysbgon 14d agoAnthropic happily paid billions to settle a lawsuit for pirating books. It's a trivial cost of doing business. If you're lucky you'll get a pittance after the fact by suing them, but a contract doesn't prevent them from doing the thing you don't want them to do and that they are obviously going to do given their past behaviour.
- nrmitchi 14d agoThe guarantee on this is a (contractual) “trust me bro”, and a right to try to sue a multi-trillion-dollar company who will absolutely drive you into the ground with legal red tape. If you are big enough to be able to withstand that, you’re already running (or trying to run) your own/open-weight models.
- torginus 14d agoJust a thought experiment: considering training seems to be 'fair use', I wonder if they trained a tiny model to retain key info from your prompts, would mean that this would still constitute fair use, and allow them to legally claim they don't retain your data.
- ronsor 14d agoZDR is shorthand for a more specified agreement of "we don't do anything other than generate your output tokens", so no. Besides, true ZDR is usually offered by third-parties with deals to host OpenAI models, such as Amazon (AWS Bedrock) and Microsoft (Azure).
- Yizahi 14d agoA lot substance is hinged on the exact definition of the word "data" or "user data". In the age of post-truth everyone is claiming that they keep no "user data". Except that after running it once through some transformer program it's no longer "user data", it's something entirely else and these corpos gave ZERO promises regarding such laundered/transformed data at all, ever.
- Aurornis 14d ago> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work. They’re not saying that some prompts are mysteriously jumping into training data. Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers.
- psyphy2 14d agoyes this is my understanding as well, and based on [1] seems to be the case. I don't know why everyone is just believing the un-backed accusations of people probably just didn't turn off said setting (and if they did why have they not said anything to such effect) [1] https://x.com/thsottiaux/status/2097746417012166816 https://x.com/thsottiaux/status/2097746417012166816
- Aurornis 14d ago> I don't know why everyone is just believing the un-backed accusations Conspiratorial thinking is very common on these topics. Even bringing up the conspiracy theory about Instagram listening to your conversations and showing you related ads will bring up a surprising amount of people defending that idea on Hacker News.
- lynndotpy 14d agoI thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs? I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.
- Gud 14d agoNo, that is not "common knowledge". You are supposed to be able to disable that unwanted feature.
- ForHackernews 14d agoI have no inside information, but I always assume the tickboxes that "disable ____ data" from Google/Facebook/OpenAI just disconnects it from your own account, not hides it from the provider.
- Aurornis 14d agoThe checkbox has an actual statement associated with it about what it does. You don’t need to assume anything.
- chii 14d agoThere's no independent verification of what that checkbox actually does. The company can say anything, and you are unable to verify that they actually do it. The only verification you could do so far is GDPR-style data export, and also the adherence to GDPR regulations (and even those might get skirted if they aren't operating in europe).
- taneq 14d agoAren’t they usually phrased very specifically as “we collect this data and use it to show you relevant ads, you can opt out of us showing you relevant ads”?
- Betelbuddy 14d agoOr these customers could just use AWS Bedrock...but their current CEO is an incompetent MBA unable to publicly articulate their biggest advantage, in the context of the current AI usage my companies. You have access to all the frontier models, but...your inputs are not shared with the model vendors...neither are used to train the next model. Why am I even doing the Amazon board job for them!??
- ballon_monkey 14d agoBedrock is really bad. It seems like they don't host the models very well because they produce tons of bugs/errors calling the model. For example you can end up with Anthropic models not returning a stop token and you end up waiting for a timeout thinking its doing something when it isn't.
- Betelbuddy 14d agoWell Anthropic hosts their models at AWS, ( and at many others...) so maybe the AWS team can ask them how they do it ;-) ?
- staticautomatic 14d agoAll except Gemini which can be rather important depending on your use case.
- Betelbuddy 14d agoYou mean the Gemini that is even behind the Chinese models?
- staticautomatic 13d agoYes, the Gemini that might be “even behind the Chinese models” but is uniquely capable of “watching” video.
- whatshisface 14d agoAmazon is deeply invested in Anthropic and would not defame them through marketing a service whose selling point was their startup's breach of contracts.
- cma 14d ago> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). 2023: "The approach also aligned with the company’s broader deployment strategy, to gradually release technologies into the world for people to get used to them. Some executives, including Altman, started to parrot the same line: OpenAI needed to get the “data flywheel” going." https://www.theatlantic.com/technology/archive/2023/11/sam-altman-open-ai-chatgpt-chaos/676050/ https://www.theatlantic.com/technology/archive/2023/11/sam-a... https://archive.is/NmO5P#selection-979.907-979.1177 https://archive.is/NmO5P#selection-979.907-979.1177 I don't think this has been a big secret.