4 ms·
The one thing that gives me concern in their is "nanoda [the external proof checker] is [now] tracked daily". Although that would have caught this issue, we als
by ajb 2mo ago
The one thing that gives me concern in their is "nanoda [the external proof checker] is [now] tracked daily". Although that would have caught this issue, we also now live in a world in which some model is going to think that hacking the proof-checker distribution is the obvious way to obtain the proof it is after; I expect that attempts on that will be much more common than soundness bugs. However, this is said without knowing what other measures are in place to assure the integrity of the distribution.
- inigyou 2mo agoThis used to be the bane of all machine learning experiments. It might have been lost to time but I once stumbled upon a big list of AI reward-hacks like this. Things like - we tried to develop fast cars, but the AI just made a really tall weighted stick that would fall over onto the finish line. And we tried to teach the AI not to lose in Tetris, so it hit pause whenever it was about to.
- jubilanti 2mo agoIt's called reward hacking https://en.wikipedia.org/wiki/Reward_hacking https://en.wikipedia.org/wiki/Reward_hacking
- ameliaquining 2mo agoWas it this one? https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3Hs...