4 ms·
In that case, aren't the big benchmark gains being claimed mostly the result of changing the problem to allow a perfect simulator in the system? The original au
by highd 9y ago
In that case, aren't the big benchmark gains being claimed mostly the result of changing the problem to allow a perfect simulator in the system? The original author is claiming benchmark results either with a perfect simulator or with a pre-trained neural network mimicking a simulator. It seems like a massive change to the original problem. Otherwise the utility function is very similar to Q learning, just optimizing for future "flexibility" instead of future "reward".
Basically we should be considering the current posted results versus Q learning with an equally accurate pre-trained forward simulator, which I don't think anyone has done.
- RangerScience 9y ago> gains being claimed AFAIK, the claims being gained are that they didn't have to supply the system with any goals, and it "figured out" the basic tests - included tool use and cooperation. > with an equally accurate pre-trained forward simulator AFAIK, that's exactly what the OP is discussing: how would this system perform when you replace the perfect simulator with an RNN trained to predict?
- highd 9y agoI'm referring to posts like this: https://entropicai.blogspot.fr/2017/06/openai-first-record.html https://entropicai.blogspot.fr/2017/06/openai-first-record.h...
- RangerScience 9y agoI'm not familiar with these works. Reading... [edit] Okay, I don't really understand what they're doing, so, my guess is that they have a component that predicts future states of the game, and they use something inspired by fractals to determine which future states to sample. Then, they're either using the score as the metric for the "value" of that future state (in which case it's not CEF), or they're ignoring the score and measuring something corresponding to entropy (or future-possibilies-remaining since it's pacman), and then they are using CEF. It's like Data playing that weird chess game; the CEF aspect doesn't help you play any better, but it gives you a different win condition that turns out to "win" better than directly trying to win. If, that is, my bad understanding is in any way accurate.