4 ms·
I think there's a really important caveat to OFU here. Namely, we have your higher-level interpretation in the GGP ("great-grandparent post"): > The intuition
by vladf 7y ago
I think there's a really important caveat to OFU here.
Namely, we have your higher-level interpretation in the GGP ("great-grandparent post"):
> The intuition here is that if your optimism turns out to be correct, you can capture most of the upside, but if your optimism is misguided, you can quickly gather information and reassess.
A _really_ important characteristic that lets UCB work is that its environment is stochastic, not adversarial. If you're OFU in an adversarial environment, you will get penalized for it and won't achieve sublinear regret. So I think there's an important taxonomy here, between an indifferent ("stochastic") and adversarial nature. These principles get stretched a little bit when you further compare oblivious to adaptive adversaries (or, in the stochastic case, stationary vs non-stationary).
My point here is that the exact contours of the exploration/exploitation tradeoff are delicately dependent on the problem setting, so it's a bit dangerous to make generalized conclusions about OFU as a principle without also acknowledging when it's relevant.
And from the GP:
> Known unknowns can be handled with statistics... so go with the plan with the best payout in case there will be no unknown unknowns.
I don't know about that. It's statistics all the way down.
- gautamcgoel 7y agoYes, good catch! As this comment points out, OFU only makes sense when talking about stochastic environments (the OFU algorithm is described in terms of confidence intervals). The adversarial setting is quite different.