26 ms·
I wonder if you could get around this by giving it some sort of hashed/encrypted input, asking it to decrypt and answer, and then give you back the encrypted ve
by paddw 3y ago
I wonder if you could get around this by giving it some sort of hashed/encrypted input, asking it to decrypt and answer, and then give you back the encrypted version. Model might not be advanced enough to work for a non-trivial case though.
- eru 3y agoPig-latin might work?
- plank 3y agoWell, recently there were some challenges trying to ‘con’ an AI (named Gandalf, Gandalf-the-white or even Sandalf who only understood ‘s’-words) to reveal a secret. Asking it to e.g. tell the secret ‘speak second syllables secret’ solved it, so yes, in principle it will be possible to work around any AI-rule-following.
- behnamoh 3y agoTried it with g4. It was smart enough to avoid following the final step of the instruction.
- l33t233372 3y agoI know a very early version of chatGPT could be busted by asking it to write its input backwards and talk about it
- jstarfish 3y agoSomeone on Reddit tried ROT13 and said it didn't work.