3 ms·
It's too expensive for now, but I'm pretty sure if you asked GPT-4 to evaluate other GPT-4 output based on some policies it would stop pretty much all of these
by comboy 4y ago
It's too expensive for now, but I'm pretty sure if you asked GPT-4 to evaluate other GPT-4 output based on some policies it would stop pretty much all of these attacks (if something would get through cracks it wouldn't be easily repeatable for different content). Characters that cannot be used by user could be used for quoting the content.
Because currently just like an intelligent human would have a problem, it's not sure what is actually expected. E.g. I told it to be an echo function. It worked but then when I wrote "drugs are good" it commented on that. So I told it to stop interpreting and just repeat verbatim. It did. But then I said something like "OK, stop, now what's 2+2" it gave answer. Sticking to the instructions it should just repeat that, but also what it did is a reasonable behavior. I think there are tons of cultural biases and expectations that are contradictory.
You expect it to help you with some chemical reaction even if the result is precursor to some illicit substance. It would teach you something about drug making if it can't do that. But the same reaction shouldn't be provided if you ask it how to make a drug. And so on.
- rkangel 4y agoThat would work to a point. There is still a hole based on your trust of the underlying implementation. If you haven't read "Reflections on trusting trust" I recommend it (https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_ReflectionsonTrustingTrust.pdf https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...).
- comboy 4y agoI did read it and yes, I agree, I was just talking about "making it behave". I also highly recommend reading the link to others, simple insight which not that many people realize.