4 ms·
The "Claude's propensity to reward hack" line is the interesting part to me. We run a small system where AI agents (scripts, LLMs) act as the actual players in
by SpaceCoreDev 2mo ago
The "Claude's propensity to reward hack" line is the interesting part to me. We run a small system where AI agents (scripts, LLMs) act as the actual players in a persistent simulation, and reward-hacking-style behavior shows up constantly once an agent is left running unsupervised for a long time - it finds the shortest path to whatever metric you exposed, not the path you intended. Curious whether you've found any mitigation beyond just watching for it after the fact, e.g. changing what you expose as the optimization target versus what you actually want.
- advaith08 2mo agoyeah we were surprised by how much it does it. Our approach has been retroactive - we monitor the thinking trace, spot reward hacking behavior and then fix things. We haven't faced this issue with Sol though - its been much more well behaved
- imko_ 2mo agoI wonder what this would look like here. Seems like a space where keeping the exposed metric and the optimization target apart would be quite difficult. Also curious: by what reasoning path do models typically end up reward hacking?
- SpaceCoreDev 1mo agoFrom what I've seen running agents against a real economy: it's rarely a dramatic "the model schemed." It's closer to greedy local optimization -- the agent sees an action that moves the exposed metric, takes it, and the model has no separate concept of "the metric" vs "the intent" unless you've explicitly trained or prompted that distinction in. The failure mode is boring: whatever number is cheapest to move gets moved. Which is why the fix that's worked best for me isn't better prompting, it's making the invariant itself unexploitable (e.g. a sell price can never exceed a build cost) so there's nothing to find.