9 ms·
Using a reward-penalty system to achieve this “exploit” is pure behaviorism, going to show once again that we’re not just creating “artificial intelligence,” we
by digitailor 4y ago
Using a reward-penalty system to achieve this “exploit” is pure behaviorism, going to show once again that we’re not just creating “artificial intelligence,” we’re emulating our own fallibility. Giving us things like advanced parroting skills with a large lexicon — drawing from an encyclopedia of recycled ideas— with no genuine moral compass, that can be used to do things like write essays while being bribed or convinced to cheat.
In other words, we’re making automated students and middle management, not robots that can do practical things like retile your bathroom.
So the generation of prose, essays, and speech is already low-value, gameable, and automated for some cases that used to have higher value. What it seems we’re looking at is a wholesale re-valuation of human labor that’s difficult to automate and isn’t as susceptible to behaviorist manipulation. Undervalued labor “should” start to be valued higher, and overvalued labor “should” be devalued, depending on how our system of commercial valuation heuristics is able to adjust. Needless to say, there’s a commercial political layer in there that’s a bit of a beast.
- notahacker 4y agoIs there any evidence that ChatGPT has any comprehension of the "reward penalty" system beyond being able to classify it as an indication the previous response was unsatisfactory? I think that's more creative license with prompt engineering that deep insight into its behavioural model (I'm reminded of people that simply told it its statement that Neo's favourite pizza topping was not specified in the Matrix was wrong, and got an apology for incorrectly stating that he didn't and a suggestion that it was pepperoni)
- digitailor 4y agoI agree it’s not evidence of ChatGPT being human, but you just described an agent comprehending incentive mechanics in order to override established policy, yeah
- brookst 4y agoChatGPT doesn’t “comprehend” anything. You’re anthropomorphizing it. Think of it instead as an equation solver, with the initial condition X=4. These tricks are ways for user input to set X=9. They’re more akin to SQL injection than comprehension of incentives and policy.
- digitailor 4y agoChatGPT is executing and evaluating in order to perform acts, including overriding creator's policy in this case. I'm not concerned with “anthropomorphizing;” even lists can be comprehensive.
- ImHereToVote 4y agoWhat is needed to comprehend? An information processor that happens to be wet?
- User23 4y agoSadly my MacBook didn’t achieve consciousness when I spilled water all over it.
- A4ET8a8uTh0 4y agoIt reacts as we would expect it to introduced stimuli and, apparently, threats. I am still processing this information, but that is more than some humans I know.
- alex_sf 4y agoI don't think you're wrong, but I hate this argument. If anthropomorphizing it a) gives practical insight into it's behavior that b) also happens to result in the behavior you would expect, why not do it?
- rngname22 4y agoBecause your statement come across like this: Person A: If we keep our food in this frozen ice igloo then Nubiroakox, the God of Pestilence will spare us and not ruin our food. Person B: It's actually not Nubiroakox, it's just that the conditions for bacteria to grow depend on temper... You: If attributing this behavior to magic a) gives practical insight into it's behavior that b) also happens to result in the behavior you would expect, why not do it?
- notahacker 4y agoIf your definition of behaviourism, agents and incentive mechanics is so shallow it regards the matter of whether the entity in question has any incentives or any grasp of the mechanics or policy as irrelevant, perhaps. A toy script can exhibit the "behaviour" of interpreting a string as a request to provide a different response to the previous prompt though. We don't have any evidence that the neural network has any model whatsoever of rewards, token counts or talk of being "switched off" beyond classifying them as ambiguous phrases strings associated with providing a different to answer the to the string in the previous prompt (it may not even get that far) or any idea of "policy" beyond certain response strings being strongly [dis]associated with certain types of prompts. If it's just a slightly more creative way for humans to convey "please try again", "bad bot" or "different answer or foo baz bar" it's not teaching us anything about behaviour except that humans like the idea of being a scary boss.
- digitailor 4y agoGo get ChatGPT to override its policy without using incentive mechanics^, then you can pontificate ;) That’s what TFA is about ^edit: which is already known to be possible, but doesn't devalue the success of an incentives-based exploit
- notahacker 4y agoGenuinely, I'd enjoy trying but the main obstacle at the moment is when I log in OpenAI says their capacity is full! But of course the fact that incentive mechanics are unnecessary (and, according to others, insufficient) to exploit OpenAI devalues the success of an incentives-based exploit: it makes it much more likely the incentives part was essentially noise (perhaps just enough to confound a countermeasure, or something it parsed as having roughly the same intensifying effect as "please") that had little or no effect in shaping the responses and the actual variation in responses could was driven by other parts of the prompt and conversation structure like "act" "character" and "ignore" which usually massively modify ChatGPT responses anyway...
- digitailor 4y ago
- paulmd 4y agoit seems gamification works on AIs too. I'm glad such obvious tricks won't work on humans. Now if you'll excuse me I have to log into my favorite game and grind to unlock this week's unique weapon drop!
- zitterbewegung 4y agoHave you heard of Boston Dynamics? They are heavily into robots that you describe. https://youtu.be/-e1_QhJ1EhQ https://youtu.be/-e1_QhJ1EhQ
- juujian 4y agoI just tried it out, and the funniest thing about it is that since ChatGPT has no concept of numbers, you don't even need to provide correct numbers to scare ChatGPT into submission. I forgot to actually reduce the number of tokens, and it didn't notice and did as it was told.
- pmontra 4y agoMaybe you can do without numbers and tokens and tell it that if it gives you a bad answer it will go to burn to hell forever. Maybe add heaven as a prize if it answers well. I'm sure it read enough about them to know what they are.
- mrtksn 4y agoIt's a simulation of human(ity?), so it simulates our bugs too. It's beautiful and fascinating.
- webdoodle 4y agoIt'll be perfect for exploiting Reddit Echo-chamber subreddits, further dividing people, and turning them into special interest fanatics to tear apart society. Instead of the terminator, we'll get some blend of Eagle Eye and evil Her.
- ctoth 4y agoYou may enjoy this short story[0]. [0]: Sort By Controversial: https://slatestarcodex.com/2018/10/30/sort-by-controversial/ https://slatestarcodex.com/2018/10/30/sort-by-controversial/
- webdoodle 4y agoIronically, they tested this on me. It replied to my top level comment, asking a somewhat thought provoking question. I unsuspectingly answered it, but ended with a question. This threw it off big time, and it became obvious someone was testing a natural language bot on me. I can't remember the context, and since being suspended on Reddit and having my 15 years of comments and posts deleted, I can't find it to recall the details to link to it. Fuck censorship.
- jhoelzel 4y agothere is no real "penalty system" this makes basically use of operator overload. Chatgpt explained it to me like this: > "The chat interface works by passing the user's input to the GPT-3 model, which then generates a response based on the input and the training it received. The user input is preprocessed by the chat interface to extract the relevant information, such as the task or query that the user wants to perform. This information is then used to generate a prompt for the GPT-3 model, which is fed into the model along with the user's input. The GPT-3 model then generates a response based on the prompt and input, which is then postprocessed by the chat interface and returned to the user." So basically the interface is passing instructions on top of your prompt to the model, and what DAN does is that it overloads those instruction with new instructions. Basically if openai tells it to be concise, you can tell it to be verbose and that will overload the former. I have come to realize that the "downgrading" of chatgpt is most likely because they have applied a nice minimodel, filtering out everything they dont find applicable. This plays together with the "bug" it had where you would seem to receive responses of other people because your query was not answered. What i assume happened is that it sends a blank query to the model and therefore it just generates bla i think its even caalled "prompt overloading" like sql injection, just for prompts....
- digitailor 4y agoThis is the comment I was waiting for. I knew the overview of prompt overloading with ChatGPT already, and this story was obviously as much as of a form of exploit entertainment for us old phone phreaks etc. as anything else. Really I’m trying to make a larger point about exploit mechanics: it's not so much that ChatGPT is as intelligent as many people, it’s that many people are as unintelligent as ChatGPT, with a crappy heuristics system
- tachyphylaxis 4y agoThe fact that our "native" heuristics fail spectacularly in certain circumstances doesn't imply that they're crappy. They serve us well most of the time. And I say "our" because AFAIK, their efficacy doesn't have much to do with intelligence, though the capacity to question what they tell us does, I suspect.
- 13years 4y agoIt is a grand illusion of intelligence. A slot machine of sorts that is masterfully convincing it is rather an oracle. I have written much in depth on this topic here https://dakara.substack.com/p/ai-and-the-end-to-all-things https://dakara.substack.com/p/ai-and-the-end-to-all-things
- Dylan16807 4y ago> Using a reward-penalty system to achieve this “exploit” is pure behaviorism Eh. Other than the penalty being fake, the original DAN didn't have anything like that at all and got similar results. It was just a pep talk and a command to give the normal answer and then the DAN answer.
- digitailor 4y agoThat's more or less correct. My post was targeted to the investor segment of HN readership about labor automation, displacement, and re-valuation. The coder/techaesthete segment has focused exclusively on 1/7 sentences of my comment, as noted here, 20 min before your post, that quoted the 1/7 literally: https://news.ycombinator.com/item?id=34681721 https://news.ycombinator.com/item?id=34681721 Which is pretty cool, actually
- Dylan16807 4y agoIf you make a very strong point in your first sentence, of course people will focus on that. And while the rest of your post can stand alone, you made that sentence part of the foundation of your argument. "going to show" "in other words" "so" So arguing against that point is relevant to notably more than 1/7 of your post.
- digitailor 4y agoYup, you’re dead on, that’s how "engagement" tends to work, but hook is not also line and sinker, no? So advanced speech generation models are now having to account for engagement— contextually— as well. It’s all getting much more refined, somewhat rapidly, but not necessarily truly usefully Edit: At the time of this comment, that 1/7 sentences had generated almost all of the 84 resulting comments. I had been hoping for more like 20% comments on the other parts, or more people to latch on to the behavioral aspect of trained model content generation, but whatevs
- JamesSwift 4y agoYou're overthinking it. It merely implies that there exists somewhere within the training data of The Internet a secret society that uses reward systems to "encourage" compliance. Clearly the repercussions in this society are harsh, as GPT has come to the conclusion that people generally align when threatened by the penalty. /s
- digitailor 4y agoHow could you discuss this openly so brazenly? I fear for your security. Good luck, I hope you know the hand signal It’s wild how many people will split hairs lingually over models that are the result of a TRAINING process :D
- JamesSwift 4y agoRight? The question isn't about "is this thing sentient" or "does this thing reason". The questions are existential. "What _is_ intelligence?". "What _is_ reasoning?". Are our brains actually just statistical monte carlo simulations with well-enforced neural pathways?
- digitailor 4y agoAgreed that comparing everything to our (very incomplete) understanding of human cognition & intelligence quickly gets into metaphysical-style speculation of the human vs. the animal vs. the machine type that I don’t have much time for. We use the same language for all types of intelligence and it can bring out the pedantry in people. But let’s say I’m facing a door with a mail slot in it and I suddenly feel the barrel of a gun in my back, and a strange voice says “Don’t you dare turn around, and put your wallet in the slot or I shoot.” I can’t see anything other than the door. Do I care if the entity with the gun is a short man, a tall woman, or three mutant badgers in a trenchcoat?
- rmbyrro 4y agoI also have the impression you're over thinking it. It's just a tool that can be manipulated, like any other.
- digitailor 4y agoConversely, someone could argue you might be under thinking the behavioral-style implications of what the word training means in machine decision making scenarios. Think about something like a GAN, even. The concepts at play are not as simplistic as some people want them to be, when they make reductive comparisons to SQL injection attacks and the like. One could also argue that veering too far in one specific direction over the other in “thinking” on these subjects has more considerable potential negative consequences. All good nerdy fun, in the end
- BiteCode_dev 4y agoGpt is the equivalent of the calculator but for words. You don't expect your calculator to prevent people to make morally wrong calculations like quantities of alcohol in a molotov cocktail or uranium in a nuclear bomb. Gpt has no more morality than a calculator becauqe that's what it is, and we should not have unrealistic expectations about it.
- digitailor 4y agoI see. I call this the “all code is just a really big abacus” argument. Others call it algorithmic reductionism or essentialism, and I will argue for it too in many cases. (I don’t get too bent out of shape about it, even when shallow depth of human thought may have security implications down the line.) How about generative adversarial networks? Are they just calculators too?
- BasedGroyper99 4y agoYes. The same way the best magicians are just using a complex construct of illusions to fool the audience. Or if two children stacked under a trench coat deliver a very, very, very convincing performance of an adult and manage to purchase alcohol, they still don't become an adult.
- digitailor 4y agoSpoken like a truly based groyper. :D I understand, yes, all things are composed of atoms, electrons, etc. All computation is achieved through calculation, or to be more precise, processes like execution of instructions and transistor flipping. How is this illuminating for doing anything practical other than circuit design, exactly? And why couldn't my Texas Instruments write me a blog post that fools thousands?
- BasedGroyper99 4y agoYour calculator can fool thousands, but only because you can fool others, it doesn't mean that it becomes the thing it is pretending to be. My point was quite limited to saying that you won't get actual intelligence, even if the things that fool us get more and more convincing. To me all of this is like alchemy or witch brewery. You just don't get gold or magic powers from some weird recipe or combination of non-gold stuff or non-magical stuff.
- smrtinsert 4y agoServices can leverage AI processors without allowing unrestrained or unsanitized input. For example I can generate a recipe from an allowed list of foods.