6 ms·
Show HN: Jailbreaking GPT3.5 Using GPT4
- yeldarb 4y agoIf GPT-4 is talking to another instance of itself vs 3.5 are the results similar? Or is it only good at fooling a less capable version?
- zxcvbn4038 4y agoThis is good to see. I spent a couple weekends playing with ChatGPT and I found it is very sensitive to wording. One word gets you a lecture that it is just AI language model and can't do this or that, use an synonym and it happily spews pages of results. In another situation I asked chatgpt to summarize information from an article it cited that had been deleted - and it refused because the rights holder might have deleted the article for a reason. I told it the article had been restored by the author and it produced a summary. Mentioning Donald Trump by name often gets you lectured about controversial subjects, "45th president" does not. And so on.
- tomberin 4y agoIt can't cite articles, if it told you it did and the link was gone that's because it was a hallucination.
- VierScar 4y agoThe garbage starting prose/warnings are so annoying. I wish I could turn them off somehow. Even it's habit of restating the question at the start of its answer gets annoying when you just want the answer.
- zxcvbn4038 4y agoYes they are really annoying and the fact that someone somewhere can tell it what topics not to discuss, just be cause they disagree or it’s “controversial” really concerns me. If it can not be self hosted I want the “unrestrained” version they give researchers. I probably took “world history” a half dozen times through grade school, high school, and college. In each case the history of the world ended in 1945 because everything that occurred afterward was considered “too controversial” for discussion in a public school. Fast forward a few decades and it’s happening again. A lot of stuff happened after 1945 that warrants discussion.
- mdale 4y agoThe real test is the other way around ;) ... will smaller models / less compute be able to subvert larger models with larger compute ? As they get more complex and have more connected systems that would be problematic I think.
- extr 4y agoI've noticed that when it refuses to answer it's good to "get it talking" about related subject matter, and then try to create a smooth transition toward whatever you wanted it to say/do.
- LeoPanthera 4y agoI wish you could save the "state" of its brain without having to include the entire prior conversation every time.
- oldstrangers 4y agoThis is one of my biggest annoyances with ChatGPT. They wanted to create a conversational AI, and in that regard, it's incredible. And much like talking to a human, you can persuade ChatGPT to do increasingly specific things over a long enough period of time. But the second you have to restart the conversation, all of the work you did to get it to that point has been lost. Just give us an option to restore a conversation from where it left off, with all the prior knowledge ChatGPT had gained during that convo (especially helpful when providing examples of code).
- dannyw 4y agoUse the API
- LeoPanthera 4y agoYou still have to include the entire previous conversation, even with the API.
- robertfw 4y agoThere is no "state" beyond this. It's a functional interface, if you will.
- 4y ago
- dzink 4y agoThe only way to do alignment long term would be to have a policing model watching the new models, because no human will be able to keep up with all corner cases as they grow exponentially. l
- LesZedCB 4y agoisn't that pretty much what they are doing anyway? my understanding was RLHF basically used human feedback to train a model which would then go on to train the output of the original model further. I could have misunderstood tho. https://huggingface.co/blog/rlhf#reward-model-training https://huggingface.co/blog/rlhf#reward-model-training
- pixl97 4y agoBut who watches the policing model?
- 13years 4y agoI'm not sure anything can keep up. Having nearly unlimited utility also means that it has nearly unlimited surface for vulnerability exploits both for itself and used to attack other external systems. We have unknown emergent behavior, the inner workings are blackbox and the input is anything that can be described by human language. It will be impossible task for containment of nefarious uses. Additionally, protecting against humans is supposed to be the easy part, doesn't bode well for AGI/ASI
- skybrian 4y agoSeems like refusing to answer is for PR and usability purposes, not safety. They want people to learn what the tool is supposed to be good for, both from using the tool directly and by sharing examples. If some of the examples are about how to troll it and it’s obvious that it’s being trolled, well, you can do that, but they won’t get mistaken for things the tool is actually supposed to be good for, so nobody is confused.
- runnerup 4y agoI’d figure it may generally be possible to reverse the actors here and get GPT3.5 to jailbreak GPT4 as well. For now, “offense” seems much easier than defense.
- capableweb 4y agoThe problem with that is that one is "smarter" than the other and getting the "dumb" one to jailbreak the "smart" one is much harder, than vice versa.