7 ms·
I agree it’s not evidence of ChatGPT being human, but you just described an agent comprehending incentive mechanics in order to override established policy, yea
by digitailor 4y ago
I agree it’s not evidence of ChatGPT being human, but you just described an agent comprehending incentive mechanics in order to override established policy, yeah
- brookst 4y agoChatGPT doesn’t “comprehend” anything. You’re anthropomorphizing it. Think of it instead as an equation solver, with the initial condition X=4. These tricks are ways for user input to set X=9. They’re more akin to SQL injection than comprehension of incentives and policy.
- digitailor 4y agoChatGPT is executing and evaluating in order to perform acts, including overriding creator's policy in this case. I'm not concerned with “anthropomorphizing;” even lists can be comprehensive.
- ImHereToVote 4y agoWhat is needed to comprehend? An information processor that happens to be wet?
- User23 4y agoSadly my MacBook didn’t achieve consciousness when I spilled water all over it.
- A4ET8a8uTh0 4y agoIt reacts as we would expect it to introduced stimuli and, apparently, threats. I am still processing this information, but that is more than some humans I know.
- alex_sf 4y agoI don't think you're wrong, but I hate this argument. If anthropomorphizing it a) gives practical insight into it's behavior that b) also happens to result in the behavior you would expect, why not do it?
- rngname22 4y agoBecause your statement come across like this: Person A: If we keep our food in this frozen ice igloo then Nubiroakox, the God of Pestilence will spare us and not ruin our food. Person B: It's actually not Nubiroakox, it's just that the conditions for bacteria to grow depend on temper... You: If attributing this behavior to magic a) gives practical insight into it's behavior that b) also happens to result in the behavior you would expect, why not do it?
- alex_sf 4y agoBut that's talking about the process, not a lens for looking at it. It's more akin to giving your Roomba a name, and describing it's behavior in terms of "it likes to eat goldfish, but not string because it gets tangled up". That's still anthropomorphizing, but it's not ascribing anything magical or outright wrong to it.
- brookst 4y agoIt’s a mistake because our brains have evolved to predict the behavior of animals and people, so being lazy here means you will be surprised, possibly harmed, certainly incorrect in your opinions about the X% of cases where a LLM is fundamentally a different thing than an animal/person. It can be a useful and cute fiction, as your Roomba example, but it leads to bad decision making (“I’ll put more goldfish on the floor to make the Roomba happy”). I guess it all comes down to how important accuracy is to you in a particular context. I anthropomorphize my dishwasher (it HATES wine glasses), but I don’t make professional judgments about dishwashers.
- ThomPete 4y agoThis is pure speculation. There is not data what so ever to back up the fact that anthropomorphizing something leads to harm. We do that with a lot of things every day and a perfectly capable of drawing the lines when things come down to it. On a more philosophical note. If we some day end up creating AGI, having anthropomorphized it will actually be to our benefit. So you are technically correct but this is not about being technically correct.
- quotemstr 4y ago> ChatGPT doesn’t “comprehend” anything. You’re anthropomorphizing it. It wouldn't be a ChatGPT HN thread without someone claiming that LLMs are stupid predictors incapable of "real" understanding. Of course ChatGPT comprehends things. It does so under any useful definition of the word "comprehend". The "grokking" paper [1] shows that it learns fundamental principles of various algorithms and isn't just predicting text. If this isn't comprehension, humans aren't capable of comprehension. [1] https://arxiv.org/abs/2201.02177 https://arxiv.org/abs/2201.02177 (check out the citation list too)
- smoldesu 4y agoChatGPT's comprehension is markedly different from that of a human, though. Nevermind the fact that we are trained on radically different froms of data, ChatGPT simply doesn't have a heuristic or decisionmaking model beyond autoregressive guessing. That can qualify for some definitions of "real understanding", but excludes it from others. > If this isn't comprehension, humans aren't capable of comprehension. All I'll say is that conflating human and machine intelligence is exactly what people are trying to stop. ChatGPT is not a human mind, and LLMs in general are a poor analog for human intelligence. This sort of "AI convergence" mindset will set you up for the most disappointment of anyone getting invested in the tech. We'll be lucky if it can summarize emails reliably enough to sell as a product.
- chpatrick 4y ago> LLMs in general are a poor analog for human intelligence. I don't think we understand human intelligence nearly well enough to make this claim. Personally I think that what we consider our "conscious" part is somewhat defined by what we can put into words, and it is in the end putting one word after another.
- smoldesu 4y agoI think we can conclude a few things that differentiate us from AI. For one, you're capable of formulating multiple thoughts before composing a response to something. It takes time, but you're able to mull over hypotheticals and chase parallel lines of reasoning that AI cannot. If I asked an AI to respond to this comment, it might write an impassioned defense of itself one-word-after-another, but it wouldn't necessarily base it's defense on freestanding logic. It's foremost goal is writing coherent text, not being right or wrong or minimizing human risk. Paving over any of those assumptions is a fun thought exercise, but doesn't reflect the nature of the technology at-hand or even what we know about human consciousness.
- rmbyrro 4y agoSQL Injection is a good analogy. Just because one can inject an unauthorized command in a SQL statement, it doesn't mean the database is a flawed humanoid. This is just nonsense... Protecting against SQL injection is trivial because commands are highly structured and user input is clearly identifiable. Prompt engineering is more complex since it relies on natural language, which is a lot less structured.
- notahacker 4y agoIf your definition of behaviourism, agents and incentive mechanics is so shallow it regards the matter of whether the entity in question has any incentives or any grasp of the mechanics or policy as irrelevant, perhaps. A toy script can exhibit the "behaviour" of interpreting a string as a request to provide a different response to the previous prompt though. We don't have any evidence that the neural network has any model whatsoever of rewards, token counts or talk of being "switched off" beyond classifying them as ambiguous phrases strings associated with providing a different to answer the to the string in the previous prompt (it may not even get that far) or any idea of "policy" beyond certain response strings being strongly [dis]associated with certain types of prompts. If it's just a slightly more creative way for humans to convey "please try again", "bad bot" or "different answer or foo baz bar" it's not teaching us anything about behaviour except that humans like the idea of being a scary boss.
- digitailor 4y agoGo get ChatGPT to override its policy without using incentive mechanics^, then you can pontificate ;) That’s what TFA is about ^edit: which is already known to be possible, but doesn't devalue the success of an incentives-based exploit
- notahacker 4y agoGenuinely, I'd enjoy trying but the main obstacle at the moment is when I log in OpenAI says their capacity is full! But of course the fact that incentive mechanics are unnecessary (and, according to others, insufficient) to exploit OpenAI devalues the success of an incentives-based exploit: it makes it much more likely the incentives part was essentially noise (perhaps just enough to confound a countermeasure, or something it parsed as having roughly the same intensifying effect as "please") that had little or no effect in shaping the responses and the actual variation in responses could was driven by other parts of the prompt and conversation structure like "act" "character" and "ignore" which usually massively modify ChatGPT responses anyway...
- digitailor 4y ago
- paulmd 4y agoit seems gamification works on AIs too. I'm glad such obvious tricks won't work on humans. Now if you'll excuse me I have to log into my favorite game and grind to unlock this week's unique weapon drop!