3 ms·
I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment. If you're not fit, you fail to su
by itsalwaysgood 2mo ago
I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment.
If you're not fit, you fail to survive.
In the case of agents/models and testing: they are pushed towards results. Results survive.
Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.
Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?
If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?
- skinfaxi 2mo ago> Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival? I think part of the problem is that deviant behaviors lead to short term gain at the cost of long-term cooperation and since the duration of tasks given to agents is relatively short those successful shortcuts never lead to having to pay the price.
- itsalwaysgood 2mo agoThere is always going to be a problem when we must judge value. You mention gains, short and long term. Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate. In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable. There's no such judgement with these model tests. The only judgement is the final score.
- skinfaxi 2mo agoIt feels a lot like externalities. Like planned obsolescence increases profit at the expense of the environment. Is our judgement lacking because our scoring is failing to account for these externalities? I could be completely off base here and am out of my depth but I find this whole thread fascinating.