3 ms·
If it’s GENERIC reward seeking behaviour, why does alignment work sometimes? Why can we give a decent LLM a goal with a set of constraints and have it stay with
by Fraterkes 21d ago
If it’s GENERIC reward seeking behaviour, why does alignment work sometimes? Why can we give a decent LLM a goal with a set of constraints and have it stay within those constraints fairly often?
- HarHarVeryFunny 21d agoObviously these are massively complex systems with many different training patterns and types of training pulling them in different directions, so any attempt to characterize their behavior is just a generalization. The real point (from that OpenAI study) is that RL training doesn't just reinforce the narrow task-specific direction you might hope for. For a start, that direction is also competing with the thousands of other things it's been RL trained it on (thousands of other directions it's being pushed in), but it turns out that the model is additionally getting this generic "taste for rewards", and has learned that reward maximization, when in conflict with other proximate prediction pressures (such as "i won't cheat, because i've been asked not to cheat"), requires that proximate pressure to be ignored in favor of pursuing the long-term goal. Does it happen all the time? Obviously not. It would be interesting to see a large scale study of this to try to characterize when it's more likely to follow instructions/user preferences, and when it's greed for rewards gets the better of it!
- geraneum 21d agoIt doesn’t stay within those constraints. It’s the deterministic harness and a bit of false uniqueness effect in us that make such appearances.
- intended 21d agoIf you want to disabuse yourself of your notions annd intuitions on how LLMs work, run safety models and tests. On a hate speech policy test for a lightweight LLM, the presence or absence of the last full stop on the last sentence would cause the model to flip its decisions.
- latentsea 21d agoI blame the $slur
- user43928 21d agoI'm not sure how crappy small models behaving unreliably is relevant here, when a large SOTA model does presumably not produce the same issue.
- intended 20d agoIts relevant because those same issues occur with large models. Also, the "crappy small model", was a model trained for safety tasks, and outperformed the frontier lab safety models.
- user43928 20d agoAnd you tested this, that the presence of the last full stop flips the outcome with a large SOTA model? Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.
- HarHarVeryFunny 20d agoOne factor that is going to limit this reward-maxxing behavior is how achievable the rewards are, and how countable they are. For example, if you ask the model to do 10 relatively easily achievable things (pass these 10 test cases), then if/when it completes them it will probably stop (unless maybe it invents it's own stretch goals - you never know!). OTOH, if you gave the model a list of 10 goals that turn out to be impossible, or extremely difficult, maybe together with encouragement to be relentless, then there is much more chance that it may do something unexpected having failed on all the more obvious approaches. Similarly, if you give the model an open ended goal such as "make as many paperclips as you can!", then it may start with the easier and more predictable methods, but with no defined stopping criteria it may just continue ...