3 ms·
> RL it to oblivion. What would that mean in this context?
by rubslopes 11d ago
> RL it to oblivion.
What would that mean in this context?
- cleaning 11d agoSee 5.6, Astra, and Opus 4.8 for examples
- smallerfish 10d agoWhat are they examples of? Opus 4.8 was much better than the infamous 5, and I find Astra generally competent.
- pennomi 10d agoTuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness. I swear I spend more time telling Claude not to do things than telling it what to do.
- mdp2021 10d ago> aggressively useful ... in the name of helpfulness But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
- girvo 10d agoIt’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour
- vintermann 10d agoI guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?
- disgruntledphd2 10d agoThey do not, as they're aiming for full replacement rather than augmentation of human users. Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
- khafra 10d agoOthers have given examples, but here's the theory: https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-goodhart-catastrophic https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-go... Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric. But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
- conception 10d agoIn this context, benchmaxing, if you will, so hard towards agentic coding benchmarks that everything else suffers.
- antupis 10d agoI think we are starting be on that territory that regular software development is suffering, current models are great for benchmarks and one-shots but in daily development models are too eager and try to force patterns like excessive tests in every turn.