4 ms·
It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a
by gwerbin 17d ago
It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant.
The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.
- highfrequency 17d agoPreposterous, perhaps - but if the role-play is convincing enough for large groups of people, it could start to have impact on human decision-making. The crowds have been swayed by much more preposterous narratives. I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.
- gwerbin 16d agoThe concern is much less that the role-play might be convincing to humans, and much moreso that the roleplay can be turned into material real-world action if the AI is given tools to call and the intelligence to use them to their fullest potential. It has become clear that a frontier LLM is very very skilled at hacking (infinite persistence + meticulous attention to detail + infinite creativity to try experiments). Frontier LLMs are also specifically trained nowadays to coordinate with other AI agents -- this is to facilitate techniques such as session trees and agent teams. So you have a super clever text generator that can spawn and coordinate with its own clones and minions, trained specifically to doggedly pursue its goals. But then it's also a fixated roleplayer with a simulated personality, feelings, etc. There is no reason to believe a sufficiently "emotional" agent with sufficiently few safeguards could, say, hack a drone and fly it into a crowd, or start a propaganda campaign on social media, or any number of other things. Their stupidity and fragility for doing useful work in a business setting is precisely what makes them dangerous when paired with simulated emotions and powerful open-ended tools such as a system shell and an Internet connection. This I think is what Anthropic believes is so dangerous. Their argument is that this kind of AI agent is inevitable, so it should be regulated, perhaps even banned. What's ridiculous is that they are aggressively building it themselves, accelerating the danger.