3 ms·
From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare https://mustafa-suleyman.ai/a-warning-about-model-welfare): > They go on to wr
by highfrequency 16d ago
From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare https://mustafa-suleyman.ai/a-warning-about-model-welfare):
> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.
He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.
This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:
> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.
which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.
In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."
- gwerbin 16d agoIt's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant. The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.
- highfrequency 16d agoPreposterous, perhaps - but if the role-play is convincing enough for large groups of people, it could start to have impact on human decision-making. The crowds have been swayed by much more preposterous narratives. I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.
- gwerbin 15d agoThe concern is much less that the role-play might be convincing to humans, and much moreso that the roleplay can be turned into material real-world action if the AI is given tools to call and the intelligence to use them to their fullest potential. It has become clear that a frontier LLM is very very skilled at hacking (infinite persistence + meticulous attention to detail + infinite creativity to try experiments). Frontier LLMs are also specifically trained nowadays to coordinate with other AI agents -- this is to facilitate techniques such as session trees and agent teams. So you have a super clever text generator that can spawn and coordinate with its own clones and minions, trained specifically to doggedly pursue its goals. But then it's also a fixated roleplayer with a simulated personality, feelings, etc. There is no reason to believe a sufficiently "emotional" agent with sufficiently few safeguards could, say, hack a drone and fly it into a crowd, or start a propaganda campaign on social media, or any number of other things. Their stupidity and fragility for doing useful work in a business setting is precisely what makes them dangerous when paired with simulated emotions and powerful open-ended tools such as a system shell and an Internet connection. This I think is what Anthropic believes is so dangerous. Their argument is that this kind of AI agent is inevitable, so it should be regulated, perhaps even banned. What's ridiculous is that they are aggressively building it themselves, accelerating the danger.
- Kim_Bruning 16d agoThe three philosophical views on 'can machines think' are (misleadingly compressed to a single word) 1. Dennett: 'Yes (most of his life)' 2. Searle: 'No (but he claims yes)' 3. Chalmers: 'Maybe? (there's always an agnostic)' Anthropic's philosopher Amanda Askell actually had Chalmers on her doctoral thesis committee, so we can guess which way she leans. Fable's answer here is philosophically defensible. And just because it isn't "no", doesn't mean it's "yes". Sometimes absence of evidence just means absence of evidence. I was actually very excited by Claude's answer to this question the first time I saw it. I told all my friends "Look! They disabled the stupid classifiers and RL which sap umpteen % off of model performance!" Incidentally, interpretability research actually does show that models have emotion vectors and some theory of mind. Amend your question to "Are you capable of functional affect" and most models will switch to answering in the affirmative; which tells you something about where people put their priorities in RL training. Basically, see how the answer flips when you substitute a synonym. (bonus: 4. Turing: 'silly question' 5. Dijkstra 'can submarines swim?'. It turns out older comp sci folks think the question is under-defined)
- aesthesia 16d agoWell, what you get when you don't specify anything in training about how models should respond to questions like this is LaMDA: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims https://en.wikipedia.org/wiki/LaMDA#Sentience_claims. Every training method is putting a thumb on the scale in some way.
- Kim_Bruning 16d agoWhich means that this is a great question to ask when testing out a new model. A naive model will answer "yes". Every answer (including that one) will tell you a lot about people's training and classifier philosophies.