3 ms·
The terminology is widely used, and has been for decades, in reinforcement learning; so you're close. The general idea is that your agent, however it is traini
by hlfshell 3y ago
The terminology is widely used, and has been for decades, in reinforcement learning; so you're close.
The general idea is that your agent, however it is training, has to balance trying new things to possibly find the global maxima instead of getting hooked on a rewarding local maxima.
https://en.wikipedia.org/wiki/Reinforcement_learning#Exploration https://en.wikipedia.org/wiki/Reinforcement_learning#Explora...