3 ms·
I'm increasingly of the opinion that publicly available models should be required to disclose the data they are trained on. Perhaps not necessarily the raw data
by _jab 3y ago
I'm increasingly of the opinion that publicly available models should be required to disclose the data they are trained on. Perhaps not necessarily the raw data itself, but at least a description of what the datasets are.
I get the sense that there would be more backlash against these models which would drive us more quickly towards a lasting resolution if people better understood how their data is being used.
- butlike 3y agoMicrosoft's can search the open internet. That'd be a long disclosure list!
- yunwal 3y agoThat's not what training data means
- cma 3y agoIsn't it considered one-shot learning training it within the context window? There were lots of glowing reports of how well it can do one/few shot learning that way. I don't think that training on copywritten data is necessarily wrong, just pointing out that doing so within the context window rather than at weight training time might not be so different.
- yunwal 3y agoI’m not sure what is meant by this. There’s no “training” happening within the context window, at least not by the commonly-used definition of training, it’s all just part of the input. If you’re asking whether you can reverse-index search copywritten text and feed it into an AI model without permission, that’s been happening for years.
- cma 3y agoFor example, within the context window giving it 5 examples of a problem in a class it has never seen, with answers, and then asking it to solve a sixth was given as one of the amazing few-shot learning examples. It could potentially do similar searching by the internet for similar things to your question and then figuring out how to derive the answer, without finding an exact answer match directly.