4 ms·
Again, we are discussing “a common pre-prompt” that you say has probability 1 of showing the system prompt… You are saying there’s some feature of this model t
by RC_ITR 3y ago
Again, we are discussing “a common pre-prompt” that you say has probability 1 of showing the system prompt…
You are saying there’s some feature of this model that deterministically returns the system prompt and then you pivot to saying that RL could never prevent something from happening.
I am saying it’s very easy to use RL to get a model to return a convincing but wrong answer about a system prompt.
Then end.
- famouswaffles 3y agoYou were wrong. Just admit it and go on with your day. This is what you said. >I am prone to believe that OpenAI, and organization who’s lead is centered on RL more than anything else, is quite good at getting it’s models not to spit out competitively sensitive information I specifically replied it is not possible to prevent a model from spitting this information out. I didn't pivot to anything. >I am saying it’s very easy to use RL to get a model to return a convincing but wrong answer about a system prompt. No it's not.
- RC_ITR 3y agoThis comment is not in the spirit of Hacker News. I was trying to co-learn by discussing with you and you turned it into something very ugly. Please do that literally anywhere else on the Internet. We clearly disagree, but I know have no idea how to move the conversation forward, which is a shame, because maybe you do have something to teach me, though I have no way of knowing at this point.