4 ms·
There's https://arxiv.org/abs/2310.00166 https://arxiv.org/abs/2310.00166, which uses an LLM for intrinsic rewards for a RL agent. They use it on nethack. It wa
by kaesve 3y ago
There's https://arxiv.org/abs/2310.00166 https://arxiv.org/abs/2310.00166, which uses an LLM for intrinsic rewards for a RL agent. They use it on nethack. It was discussed on the TalkRL podcast: https://www.talkrl.com/episodes/pierluca-doro-and-martin-klissarov https://www.talkrl.com/episodes/pierluca-doro-and-martin-kli...
- bubblyworld 3y agoOoh, really cool! Thanks for the links.