5 ms·
Slightly tangential, is there some kind of crowdsourced effort to build training data for fine tuning? Alpaca used the training data built from gpt-3.5, so ther
by Hedepig 4y ago
Slightly tangential, is there some kind of crowdsourced effort to build training data for fine tuning? Alpaca used the training data built from gpt-3.5, so there are terms of use restrictions
- Phemist 4y agoIANAL, but wouldnt feeding the training data into alpaca and then having it output similar text not construe new data that is not copyrighted by stanford/openai/facebook? (It would be a significantly creative and novel worm to get the prompts working correctly..) Obviously you are also not bound by the openai terms of use, and Im not sure if stanford's terms of use are as broad and well-defined...
- Hedepig 4y agoThe terms of ChatGPT/GPT-3.5 explicitly state that one cannot use their data to construct competitive models
- NavinF 4y agoNot enforceable. All they can do is ban your ChatGPT account.
- YetAnotherNick 4y ago> Not enforcable Are you willing to assign a upper limit on this probability and bet for it?
- throwaway1851 4y agoOpenAI disclaims ownership interest in the model output. If a subscriber (who has a contractual relationship with OpenAI) chooses to generate outputs that could be used to train a competing model, and chooses to share those outputs with third parties, that is not prohibited by the agreement. Further, the data being shared belongs to the subscriber and can be licensed however they desire (though actually, model outputs may not be copyrightable at all). If a third party who does not have a contractual relationship with OpenAI chooses to take this data and train a competing model, they are using the data under valid license and have breached no obligation to OpenAI.
- satvikpendem 4y agoDo you have the money to enforce such (dis)claims? That's really what it comes down to, you might be right but you'd go bankrupt in the process. It's simply better to just use something that didn't touch their outputs in the first place, like OpenAssistant.
- amrb 4y agoI'm not a lawyer but if a successful company was built off chatgpt's output and without having a contract in place. I could see the totally morale megaCorp's trying to legally take an ownership stake. Even recently US copywrite office have asked you too list any parts built with AI as we don't have the laws in place to cover this: https://www.copyright.gov/ai/ https://www.copyright.gov/ai/
- Phemist 4y agoYes, but 1) you would not be using their data, and 2) you are not bound by their ToS if you never signed up to their service right..
- Metus 4y agohttps://open-assistant.io https://open-assistant.io
- Hedepig 4y agoThis looks good. Is the training data in the repo itself?
- amrb 4y agoYou can get datasets here, depending on the training you want to do: https://huggingface.co/datasets/ https://huggingface.co/datasets/
- lxe 4y agoOpenChatKit is great tbh