4 ms·
Do OpenIA et al. even understand it themselves?
by thibautg 3y ago
Do OpenIA et al. even understand it themselves?
- simonw 3y agoSomething they definitely understand is exactly what their pre-training data looks like - the raw text that goes into the initial runs of training the models. Instruction tuning and RLHF is a bit more complex than that. I assume they maintain detailed logs all of those human-driven decisions about which responses were better.
- thibautg 3y agoDo they really know exactly what the raw text looks like? It seems so huge that no human could read all the text from each corpus. And the models have been fed text from many different languages. I doubt they have people who understand all the languages. Regarding RLHF, I also hope that they kept the logs of all the human decisions. But since it was (at least partially) outsourced to African companies like Sama.com, do they really get back all the logs or just a new fine-tuned model? But they must indeed at least know what is done with the text submitted to the prompt or to their API. (I’m really not an expert so my questions may sound naive)
- dontupvoteme 3y agoRLHF apparently manages to mostly force all responses in all languages to be quasihomogeneous. I'm not sure if that means they translated the RLHF data to as many languages as possible and then repeated it or if it's something more fundamental which applies regardless of input language. Although asking it "What can you not talk about" in Japanese only responds correctly with gptv4, and each language gives you a different list of items to some degree (between 4 and 6 items i found). Sadly trying to speak Klingon or Sindarin to it is dodgy at best
- fantyoon 3y agoIts safe to assume that whoever OpenAI outsources to does not get access to the model. Collecting data and training models on it will be two different steps.
- pixl97 3y ago>Do they really know exactly what the raw text looks like? I mean yea, it's too big. That said, in a post ad hoc fashion they do. When the model spits out weird crap at times, they can search the raw corpus and filter those strings out. There was an incident around this with Reddit counting forums and strange usernames that were added in tokenization, but later removed from weights leading to odd behavior when doing inference.
- pxoe 3y ago[dead]
- sebzim4500 3y agoOpenAI obviously don't know all the data in their dataset, but they must at least know whether they are training on private data submitted to their API.
- closewith 3y agoMaybe we need an Open Internal Affairs to police OpenAI?
- jjulius 3y agoSome OpenIA for OpenAI.
- rvz 3y agoI don’t think that they even understand what is going on inside of these AI models. Why is that? GPT-4 is as transparent as a black box and lacks transparent explainability and openly admits to regurgitating nonsense. How do they allegedly ‘fix’ it? Train it on more data.