4 ms·
But you don't need "most of the researchers" or relocating them, nor you need legal data - with sufficient money and hardware, if North Korea or any other count
by PeterisP 3y ago
But you don't need "most of the researchers" or relocating them, nor you need legal data - with sufficient money and hardware, if North Korea or any other country puts a bunch of median CS grad-students to the task, they can train whatever models they want; the code and data is practically available and will be, no matter what the bans.
- ben_w 3y ago> with sufficient money and hardware Is making an assumption that hardware is not also facing restrictions. > code and data is practically available Disagree; although the base model data is, the RLHF data isn't. This is also why 3rd parties are not able to replicate Google, despite all the PageRank patents having expired and the web being about as crawlable for Google as for, say, Bing. > "median CS grad-student" … is how I regard the quality of ChatGPT's code output, FWIW.
- skissane 3y ago> the RLHF data isn't There are open sources of RLHF data - for example, the LMSys Chatbot Arena dataset (for same prompt two different responses along with which the human preferred), ShareGPT It also isn’t hard to use existing leading AIs like GPT-4 or Claude-3 to generate synthetic RLHF datasets. They’ll put words in their terms of service to say you can’t use a dataset generated in that way to a train a competing model, but their ability to enforce those terms in practice is very open to question