3 ms·
> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. Build a better simula
by joshheitzman 13d ago
> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse.
Build a better simulator to train them in (i.e. more expensive) that includes a simulation of an intranet and the internet and is air gapped so there is no escape. Sneaker transfer the total system data at each step to another air gapped system to evaluate it and sneaker transfer the reward back. That the reward function has to penalize all modifications to state that are out of bounds.
Yeah, I realize that will be amazingly slow.