3 ms·
If this is RSI then all RL is RSI. The whole idea behind RL is that the agent improves over time by making actions in the environment and observing the next st
by mohamedmohey 18d ago
If this is RSI then all RL is RSI.
The whole idea behind RL is that the agent improves over time by making actions in the environment and observing the next state and the reward then modifying its policy.
The exploration-exploitation dilemma still stands. "dreaming in the replay simulator" is jargon for the same replay ideas presented when Deep RL was first introduced (Q-learning with experience replay and all that).