4 ms·
> It's not going to happen accidentally. https://transformer-circuits.pub/2026/emotions/index.html https://transformer-circuits.pub/2026/emotions/index.html W
by Philpax 18d ago
> It's not going to happen accidentally.
https://transformer-circuits.pub/2026/emotions/index.html https://transformer-circuits.pub/2026/emotions/index.html
Whether these are like "our" emotions is hard to say. What we _can_ say is that they are emotion-shaped, we didn't design them, and they happened accidentally.
Modern AI is grown, not meticulously designed, and we cannot say with any certainty what the resulting mechanistic properties are.
- HarHarVeryFunny 18d agoAn LLM will learn anything that helps it predict, including the emotional state of the writer - that is expected. If you give an LLM the move sequence of a half-played chess game and ask it to continue as white or black, then it has learnt enough to model the ELO rating of both players and will continue playing at that level. It is not playing to win - it is doing what you expect and predicting as well as it can - it predicts the 1500 ELO player will keep playing at that level, and generates moves accordingly. An LLM appearing to exhibit an emotion (if we anthropomorphize it and read emotion into it's output) is just predicting as well as it can - if the context calls for sad output, they you'd expect to get sad output and will necessarily find that "we're predicting sadness" detector somewhere internally. Transformers are the same as they ever were from 10 years ago, other than minor efficiency tweaks like MOE and different attention mechanisms. Training is getting more and more complex, resulting in better and better cargo cult reasoning etc, but the architecture remains the same.
- famouswaffles 18d ago>An LLM will learn anything that helps it predict I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow.
- HarHarVeryFunny 17d agoAre you imagining that an LLM tasked with predicting a game continuation is going to play to win instead?
- famouswaffles 17d agoI imagine it will learn to win under some circumstances, perhaps in a case with some context expressing a desire to win. Drawing out an LLMs upper ability in the game should be fairly straightforward.
- HarHarVeryFunny 17d agoIf you asked it to try to win, to "plan lines step by step", etc, then it would do it's best to follow that instruction, but unless RLVR trained to reason about chess (easy to do, but not sure which models may have done it) then it'd have to instead rely on the chess reasoning it had seen during pre-training (post-game interviews etc), which I doubt is enough to do very well. However, if you just ask it to continue a game, halfway in progress, then by default it will try to predict the most likely continuation, which is that both players will continue to play at the level they have done so far. This isn't a theory - it's been documented, as well as what you'd expect.
- famouswaffles 17d agoI mean sure, but I'm not sure what that has to do with the broader point. It will learn to play, and it will have a model of what it means to win.
- HarHarVeryFunny 17d ago> I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow I was just explaining how this comment you made is wrong.
- famouswaffles 17d agoIt's not wrong. You admit that LLMs will 'learn anything that helps them predict' and fail to realize the breadth of that statement. Your chess statements don't really help your case. It doesn't matter that it usually doesn't primarily care about winning. It still learnt how to play the game, and it still knows how to win. Similar outcomes for predicting emotions would mean it still developed an affective state, and that its ability to 'feel angry' is no less real.