4 ms·
Hmm, perhaps something like UCB1? The UCB (upper-confidence-bound) family of algorithms might be appropriate—it's not quite what you've described though. They'
by dhotson 8y ago
Hmm, perhaps something like UCB1? The UCB (upper-confidence-bound) family of algorithms might be appropriate—it's not quite what you've described though.
They'll explore the arm with the highest potential payoff i.e. the highest upper-confidence-bound, which often is the arm you know the least about since they're roughly calculated as (average + confidence interval) and early on the confidence bound is large. This style of algorithm means you can add arms as you go through the experiment and they'll be explored/exploited in a reasonable way.
What you're describing sounds more like you're exploring a frontier and "discovering" new options along the way..?