4 ms·
I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out c
by RC_ITR 3y ago
I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information.
Can you get yours to give you the same verbatim?
- famouswaffles 3y ago>I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information. Thanks for telling me you don't know how RL or LLMs work. >Can you get yours to give you the same verbatim? Sure I can. and others in this very thread have too. https://news.ycombinator.com/item?id=37805492 https://news.ycombinator.com/item?id=37805492
- RC_ITR 3y agoOk then explain why RL can’t be used to prevent certain behaviors please. Why can’t a reward function be used to stop a model from saying something you know you don’t want it to say? Also you share a screenshot of a chat asking to repeat the above and that’s your proof? Share the raw link please.
- famouswaffles 3y ago>Ok then explain why RL can’t be used to prevent certain behaviors please. Preventing certain behaviors does not mean you can make a model never output something. RL simply just doesn't work that way. In this instance, You are rating certain responses better and asking the model to predict like that. You can make it more likely to refuse a request but the idea that you can guarantee it won't is completely wrong. There is nothing open ai can do to make GPT-4 never do something. Nothing. https://chat.openai.com/share/b7faf20c-b295-4d76-85a1-a15e0481dd21 https://chat.openai.com/share/b7faf20c-b295-4d76-85a1-a15e04...
- geraneum 3y ago> You can make it more likely to refuse a request but the idea that you can guarantee it won't is completely wrong. Why is that the case, technically?
- Jensson 3y agoBecause it is a black box, they don't know enough about it to ensure it never does something. Only way to be sure is to write some script using normal code that filters the questions and outputs, but then you have the standard natural language problem which only works for very simple cases.
- geraneum 3y ago> to ensure it never does something But there are many systems for which you cannot predict/control the behavior with just a few experiments because they are simply, probabilistic. Isn’t it also the case with LLMs? If not, why?
- RC_ITR 3y agoAgain, we are discussing “a common pre-prompt” that you say has probability 1 of showing the system prompt… You are saying there’s some feature of this model that deterministically returns the system prompt and then you pivot to saying that RL could never prevent something from happening. I am saying it’s very easy to use RL to get a model to return a convincing but wrong answer about a system prompt. Then end.
- famouswaffles 3y agoYou were wrong. Just admit it and go on with your day. This is what you said. >I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information I specifically replied it is not possible to prevent a model from spitting this information out. I didn't pivot to anything. >I am saying it’s very easy to use RL to get a model to return a convincing but wrong answer about a system prompt. No it's not.