3 ms·
Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness. I swear I spend more ti
by pennomi 8d ago
Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.
I swear I spend more time telling Claude not to do things than telling it what to do.
- mdp2021 8d ago> aggressively useful ... in the name of helpfulness But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
- girvo 8d agoIt’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour
- vintermann 8d agoI guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?
- disgruntledphd2 8d agoThey do not, as they're aiming for full replacement rather than augmentation of human users. Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.