4 ms·
Interesting story, if maybe a bit oversold. Since the game apparently announces this (didn’t know that), shouldn’t the model have detected a significant differe
by jsjohnst 2y ago
Interesting story, if maybe a bit oversold. Since the game apparently announces this (didn’t know that), shouldn’t the model have detected a significant difference during the gameplay?
- elijaht 2y agoThe article mentions that “It simply doesn't have data about full moon variables in its training data, so a branching series of decisions likely leads to lesser outcomes, or just confusion.”
- thaumasiotes 2y agoIt does say that, but I'd like to know more about the mechanisms involved. The problem is that every action you take generates better results than you'd expect. I can see why prediction would be worse while that's going on, but why would performance be reduced?
- bravetraveler 2y agoI believe games like NetHack have monkey paw situations where fortune may not actually be that fortunate My memory has not lasted well
- andersource 2y agoI can imagine it's a bit like if gravity suddenly changed to 0.9g. Everything(?) is easier but a lot of people would probably stumble around a bit before muscle memory, coordination etc. get used to the change.
- seanhunter 2y agoNethack is weird. It's literally a game comprised of a huge number of special cases all welded together. A lot of the crucial info for playing the game is not really discoverable through the interface, so Humans used to mostly learn the game by reading the source code and now they mostly learn the game by reading guides. The bug here is not in nethack but just in the training, which meant some of the special cases (full moons, fri 13ths etc) weren't in the training data. They should have been running the training in VMs with the clock set to include these cases. Honestly a lot of the reporting of this "bug" seems wildly overblown.
- WJW 2y agoWell if it was only trained on non-full moon days, even if the model did detect a difference it would have no idea how to adapt its play style.
- jsjohnst 2y agoAs another commenter said, it’s quite obvious about it. For such a key difference, I’d think they’d have the model run record a log for later inspection. Or they’d be watching the game play. Or something similar.
- WJW 2y agoIf the model was trained over (say) a week during there was no full moon, how would the model know what this unexpected message means? It would probably just ignore it and continue playing as normal. Nethack is full of messages that can be safely ignored, so just another one would not be unusual. I don't agree that the creators would be watching the game play either. Usually during such training phases you'd run as many copies of the game as the available hardware can manage. I wouldn't be surprised if they had at least hundreds of runs going in parallel and the researcher is definitely not going to be watching them all. If anything, they are going to bed and train the model overnight as much as possible.
- jsjohnst 2y ago> how would the model know what this unexpected message means? It would probably just ignore it and continue playing as normal. Then that’s a poor model then. Significant anomalies should be flagged for manual review, otherwise you corrupt datasets unintentionally.
- thaumasiotes 2y agoAccording to the article, the effect is that actions have better outcomes than they otherwise would. Assuming you don't adapt your playstyle in any way, how would that lead to worse overall performance?