18 ms·
I think this is different in that ChatGPT is expressly using your data as training in a probabilistic model. This means: * Their contractors can (and do!) see
by jackson1442 4y ago
I think this is different in that ChatGPT is expressly using your data as training in a probabilistic model. This means:
* Their contractors can (and do!) see your chat data to tune the model
* If the model is trained on your confidential data, it may start returning this data to other users (as we've seen with Github Copilot regurgitating licensed software)
* The site even _tells you_ not to put confidential data in for these reasons.
Until OpenAI makes a version that you can stick on a server in your own datacenter, I wouldn't trust it with anything confidential.
- wodenokoto 4y agoWell, you can stick it on Azure.
- hgsgm 4y agoGoogle had all the same problems, until it found a balance of functionality, security, and privacy. OpenAI just hasn't started to try adding privacy and security yet.
- endisneigh 4y agoA language model inherently has a privacy problem. How would you guarantee no leaks?
- pama 4y agoYou simply don’t train on the user imputs. There are enough unread books, public repos, and new articles.
- renewiltord 4y agoNot that I don't expect them to do this, but how is it expressly said to be so? https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance https://help.openai.com/en/articles/5722486-how-your-data-is... > OpenAI does not use data submitted by customers via our API to train OpenAI models or improve OpenAI’s service offering. In order to support the continuous improvement of our models, you can fill out this form to opt-in to share your data with us. Sharing your data with us not only helps our models become more accurate and better at solving your specific problem, it also helps improve their general capabilities and safety.
- kristofferR 4y agoDid you read the next paragraph? > When you use our non-API consumer services ChatGPT or DALL-E, we may use the data you provide us to improve our models.
- renewiltord 4y agoI definitely did not correctly read that. Thanks for the clarification. Totally misread the 'our API' bit! It's also in the FAQ: https://help.openai.com/en/articles/6783457-chatgpt-general-faq https://help.openai.com/en/articles/6783457-chatgpt-general-... > Will you use my conversations for training? > Yes. Your conversations may be reviewed by our AI trainers to improve our systems.
- muzani 4y agoAs the saying goes, if the product is free, then you are the product.
- renewiltord 4y agoThe product is $20!
- avereveard 4y agoHehe old tos trick. Here it doesn't say "will never use" but say "does not use" and I wager below or somewhere will say that they can change the tos at any time in the future unilaterally
- waboremo 4y agoSticking it in your own datacenter doesn't really prevent any of these problems (except maybe #2), only now your leaks are internal and because of all the false sense of security, you might wind up leaking far more confidential and specific information (ie. an executive leaking to the rest of the team in advance that they are planning layoffs for noted reasons, whereas that executive might have used more vague terms when speaking to public chatGPT).
- mholm 4y agoSticking it in your own private datacenter would imply that you can opt in or out of using your data to train the next generation. ChatGPT does not dynamically train itself in realtime.
- waboremo 4y agoThe implication is that you would bother with ChatGPT at all to train it on the relevant local data, the key value aspect to ChatGPT beyond general public use.
- FartyMcFarter 4y agoIt prevents all of those problems as it puts all the data / data movement under your control.
- waboremo 4y agoHow so?
- FartyMcFarter 4y agoBecause it's your own data center, which means you own the data and can set up firewall rules to prevent any software running there from leaking data to outside the data center.
- jacquesm 4y ago> I think this is different in that ChatGPT is expressly using your data as training in a probabilistic model. Google tries hard to sell you on their auto-answers for emails ('smart reply'), wonder how those got trained...