5 ms·
Question about privacy: OpenAI doesn't use API calls to train their models. But do they or Microsoft still store the text? If so, for how long? Overall, I thin
by tuckerconnelly 3y ago
Question about privacy: OpenAI doesn't use API calls to train their models. But do they or Microsoft still store the text? If so, for how long?
Overall, I think this is great, and can't wait for the 16k fine-tuning.
- haldujai 3y agoNot sure about direct OpenAI API calls but with the Azure offering they store prompts and output for 30 days to monitor for abuse. There is an application form if one wants to be exempted from this requirement. https://learn.microsoft.com/en-us/legal/cognitive-services/openai/data-privacy#how-does-the-azure-openai-service-process-data https://learn.microsoft.com/en-us/legal/cognitive-services/o...
- 3abiton 3y agoDoes the finetuned model reside on OpenAI's servers? If so, what privacy guarantees that openai won't utilize it later for expanding gpt5?
- flangola7 3y agoInsist on such guarantees in the contact.
- jakeduth 3y agoYes they are stored on OpenAI's servers. The API calls are not used for model training per the TOS. However, not that I'm accusing OpenAI of anything, but there's no way to independently validate this. But their guarantee is clear for the API (the ChatGPT web app is different, but you can disable training if you give up the history feature). > At OpenAI, protecting user data is fundamental to our mission. We do not train our models on inputs and outputs through our API. > ... > We do not train on any user data or metadata submitted through any of our APIs, unless you as a user explicitly opt in. > ... > Models deployed to the API are statically versioned: they are not retrained or updated in real-time with API requests. > Your API inputs and outputs do not become part of the training data unless you explicitly opt in. - https://openai.com/api-data-privacy https://openai.com/api-data-privacy
- zarzavat 3y agoIt’s in principle possible to detect if a model has been trained on private data, e.g. if it can recite random data such as UUIDs that are not public. So if OpenAI were to break that promise, someone would notice and make it public. This is enough of a disincentive that I trust OpenAI will not do it.
- tedsanders 3y ago30 days maximum, in most cases: https://platform.openai.com/docs/models/default-usage-policies-by-endpoint https://platform.openai.com/docs/models/default-usage-polici... We don’t do anything sneaky with the stored data; literally the only purpose is to be able to investigate possible trust and safety violations for a brief period after they occur.