4 ms·
Everyone understands that these are machines that make convincing answers to questions without following any symbolic rules Until The convincing answer is som
by RC_ITR 3y ago
Everyone understands that these are machines that make convincing answers to questions without following any symbolic rules
Until
The convincing answer is something you want to believe follows symbolic rules.
Posts like these really foreshadow how valuable “knowing when to take the LLM at face value” will be as a job skill.
- famouswaffles 3y agoYou do realize this "list of rules as a pre-prompt" is common and happens right ? This isn't some hallucination (which is easily tested by asking again on a fresh instance and seeing if it's consistent).
- RC_ITR 3y agoI am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information. Can you get yours to give you the same verbatim?
- famouswaffles 3y ago>I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information. Thanks for telling me you don't know how RL or LLMs work. >Can you get yours to give you the same verbatim? Sure I can. and others in this very thread have too. https://news.ycombinator.com/item?id=37805492 https://news.ycombinator.com/item?id=37805492
- RC_ITR 3y agoOk then explain why RL can’t be used to prevent certain behaviors please. Why can’t a reward function be used to stop a model from saying something you know you don’t want it to say? Also you share a screenshot of a chat asking to repeat the above and that’s your proof? Share the raw link please.
- famouswaffles 3y ago>Ok then explain why RL can’t be used to prevent certain behaviors please. Preventing certain behaviors does not mean you can make a model never output something. RL simply just doesn't work that way. In this instance, You are rating certain responses better and asking the model to predict like that. You can make it more likely to refuse a request but the idea that you can guarantee it won't is completely wrong. There is nothing open ai can do to make GPT-4 never do something. Nothing. https://chat.openai.com/share/b7faf20c-b295-4d76-85a1-a15e0481dd21 https://chat.openai.com/share/b7faf20c-b295-4d76-85a1-a15e04...
- geraneum 3y ago> You can make it more likely to refuse a request but the idea that you can guarantee it won't is completely wrong. Why is that the case, technically?
- Jensson 3y agoBecause it is a black box, they don't know enough about it to ensure it never does something. Only way to be sure is to write some script using normal code that filters the questions and outputs, but then you have the standard natural language problem which only works for very simple cases.
- geraneum 3y ago> to ensure it never does something But there are many systems for which you cannot predict/control the behavior with just a few experiments because they are simply, probabilistic. Isn’t it also the case with LLMs? If not, why?
- RC_ITR 3y agoAgain, we are discussing “a common pre-prompt” that you say has probability 1 of showing the system prompt… You are saying there’s some feature of this model that deterministically returns the system prompt and then you pivot to saying that RL could never prevent something from happening. I am saying it’s very easy to use RL to get a model to return a convincing but wrong answer about a system prompt. Then end.