3 ms·
There is no way this could be reasonably framed as RSI. This iterative, online optimization of an exploration policy is not recursively intelligent in any way.
by bob1029 18d ago
There is no way this could be reasonably framed as RSI.
This iterative, online optimization of an exploration policy is not recursively intelligent in any way. It simply reallocates the available computational resources to more promising (hopefully) parts of the search space as system conditions change over time.
- mohamedmohey 17d agoIf this is RSI then all RL is RSI. The whole idea behind RL is that the agent improves over time by making actions in the environment and observing the next state and the reward then modifying its policy. The exploration-exploitation dilemma still stands. "dreaming in the replay simulator" is jargon for the same replay ideas presented when Deep RL was first introduced (Q-learning with experience replay and all that).