3 ms·
Here is one good overview: http://www.cs.nyu.edu/~mohri/pub/bandit.pdf http://www.cs.nyu.edu/~mohri/pub/bandit.pdf "[...] ε-greedy is probably the simplest a
by svedlin 15y ago
Here is one good overview:
http://www.cs.nyu.edu/~mohri/pub/bandit.pdf http://www.cs.nyu.edu/~mohri/pub/bandit.pdf
"[...] ε-greedy is probably the simplest and the most widely used strategy to solve the bandit problem and was first described by Watkins [24]. The ε-greedy strategy consists of choosing a random lever with ε-frequency, and otherwise choosing the lever with the highest estimated mean, the estimation being based on the rewards observed thus far. ε must be in the open interval (0, 1) and its choice is left to the user. Methods that imply a binary distinction between exploitation (the greedy choice) and exploration (uniform probability over a set of levers) are known as semi-uniform methods."