3 ms·
In the results graphs it looks like it gets beat pretty bad by the “dreamer” baseline, am I missing something?
by tehsauce 4y ago
In the results graphs it looks like it gets beat pretty bad by the “dreamer” baseline, am I missing something?
- fxtentacle 4y agoThis one can learn long-term strategies such as "stand upright so that you can walk around" without any demonstration. Dreamer works equally well, but needs a more elaborate training setup and it's not hierarchical, meaning the AI cannot learn to self-impose intermediate goals. So the novelty is that they make a hierarchical long-term RL AI trainable end-to-end, which was impossible before.