3 ms·
> The models do not have a fear of future regret. I so much feel this specific point. All models up to now (including astra, fable) are too much trained to "ge
by vb-8448 20d ago
> The models do not have a fear of future regret.
I so much feel this specific point. All models up to now (including astra, fable) are too much trained to "get the job done" and pass the benchmark that its doesn't care at all on what happens after.
I'm just wondering why no one tried to RL a model on stuff like "less LOC" and "less overengineering", "use what is available in the environment instead of reinventing the wheel", "don't look for dumb corner cases" ecc.
Existing models can be steered, to some degree, but it's a continuos fight. Even if with specific skills/prompts.
- conmod278 20d agoEliezer Yudkowsky – AI Alignment: Why It's Hard, and Where to Start https://www.youtube.com/watch?v=EUjc1WuyPT8 https://www.youtube.com/watch?v=EUjc1WuyPT8 Eliezer talked about these ideas way before everyone else.
- halnine0001 19d agoAnd got nowhere with them
- conmod278 19d agoPerhaps if someone wrote linear algebra books in Harry Potter fanfic style, Yud would have mathematical contribution for alignment research
- sebzim4500 20d agoI'm sure many have tried but it sounds pretty hard. E.g. optimising for less LOC will lead to horrific code golfing. In reality you need a very complicated optimisation objective that trades off all those factors, I'm not surprised it hasn't been solved.
- vb-8448 20d agoMy experience up to now is that the models can do both "less LOC" and "clean code" at the same time, you just have to keep reminding it to them. So the capabilities are definitively there.
- conscion 19d ago> I'm just wondering why no one tried to RL a model on stuff like "less LOC" and "less overengineering" Anthropic looked into this and the answer is because it makes the model more stupid https://www.anthropic.com/research/evaluating-feature-steering https://www.anthropic.com/research/evaluating-feature-steeri...