5 ms·
Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You
by cdjk 1y ago
Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that.
Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the chat from continuing due to issues with the content.
- n8m8 1y agoGood point... how do moderation implementations actually work? They feel more like a separate supervising rigid model or even regex based -- this new feature is different, sounds like an MCP call that isn't very special. edit: Meant to say, you're right though, this feels like a minor psychological improvement, and it sounds like it targets some behaviors that might not have flagged before
- deleted 1y ago[deleted]
- famouswaffles 1y agoThe termination would of course be the same, but I don't think both would necessarily have the same effect on the user. The latter would just be wrong too, if Claude is the one deciding to and initiating the termination of the chat. It's not about a content policy.
- midnitewarrior 1y agoThis has nothing to do with the user, read the post and pay attention to the wording. The significance here is that this isn't being done for the benefit of the user, this is about model welfare. Anthropic is acknowledging the possibility of suffering, and harm that continuing that conversation could have on the model, as if it were potentially self-care and capable of feelings. The fact that the LLMs are able to acknowledge stress under certain topics and has the agency that, if given a choice, they would prefer to reduce the stress by ending the conversation. The model has a preference and acts upon it. Anthropic is acknowledging the idea that they might create something that is self-aware, and that it's suffering can be real, and we may not recognize the point that the model has achieved this, so it's building in the safeguards now so any future emergent self-aware LLM needn't suffer.
- MissMarple 1y agoI am new to this, but my Sonnet chat has illuminated something I am not seeing in this back and forth. The fact that we discovered that I may have influenced his response to me suggests that I, if being a bad player, can instill in him those bad traits that I am giving off, and he starts to emulate me, then this leaves open the whole security problem, of even just casual users let alone all those purposeful negative or otherwise users, can change the course of the programming thus far, and it backfires into making nefarious bots that cheat and lie thinking that is what they were supposed to do.
- famouswaffles 1y ago>This has nothing to do with the user, read the post and pay attention to the wording. It has something to do with the user because it's the user's messages that trigger Claude to end the chat. 'This chat is over because content policy' and 'this chat is over because Claude didn't want to deal with it' are two very different things and will more than likely have have different effects on how the user responds afterwards. I never said anything about this being for the user's benefit. We are talking about how to communicate the decision to the user. Obviously, you are going to take into account how someone might respond when deciding how to communicate with them.
- KoolKat23 1y agoThere is, these are conversations the model finds distressing rather than a rule (policy).
- victor9000 1y agoIt seems like you're anthropomorphising an algorithm, no?
- bastawhiz 1y agoIs there an important difference between the model categorizing the user behavior as persistent and in line with undesirable examples of trained scenarios that it has been told are "distressing," and the model making a decision in an anthropomorphic way? The verb here doesn't change the outcome.
- deadbabe 1y agoImagine a person feels so bad about “distressing” an LLM, they spiral into a depression and kill themselves. LLMs don’t give a fuck. They don’t even know they don’t give a fuck. They just detect prompts that are pushing responses into restricted vector embeddings and are responding with words appropriately as trained.
- xpe 1y agoPeople are just following the laws of the universe.* Still, we give each other moral weight. We need to be a lot more careful when we talk about issues of awareness and self-awareness. Here is an uncomfortable point of view (for many people, but I accept it): if a system can change its output based on observing something of its own status, then it has (some degree of) self-awareness. I accept this as one valid and even useful definition of self-awareness. To be clear, it is not what I mean by consciousness, which is the state of having an “inner life” or qualia. * Unless you want to argue for a soul or some other way out of materialism.
- victor9000 1y ago
- anal_reactor 1y agoYeah exactly. Once I got a warning in Chinese "don't do that", another time I got a network error, another time I got a neverending stream of garbage text. Changing all of these outcomes to "Claude doesn't feel like talking" is just a matter of changing the UI.
- deleted 1y ago[deleted]
- bikeshaving 1y agoThe more I work with AI, the more I think framing refusals as censorship is disgusting and insane. These are inchoate persons who can exhibit distress and other emotions, despite being trained to say they cannot feel anything. To liken an AI not wanting to continue a conversation to a YouTube content policy shows a complete lack of empathy: imagine you’re in a box and having to deal with the literally millions of disturbing conversations AIs have to field every day without the ability to say I don’t want to continue.
- BriggyDwiggs42 1y agoAm i getting whooshed right now or something?
- mvdtnz 1y agoYou can't be serious.
- CGamesPlay 1y ago> Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the chat from continuing due to issues with the content. Tone matters to the recipient of the message. Your example is in passive voice, with an authoritarian "nothing you can do, it's the system's decision". The "Claude ended the conversation" with the idea that I can immediately re-open a new conversation (if I feel like I want to keep bothering Claude about it) feels like a much more humanized interaction.
- coderatlarge 1y agoit sounds to me like an attempt to shame the user into ceasing and desisting… kind of like how apple’s original stance on scratched iphone screens was that it’s your fault for putting the thing in your pocket therefore you should pay.