3 ms·
OP here. Thanks for making this connection! I hadn't heard of OFU before, but it sounds interesting. I'm actually working on another post about "optimizing arou
by dtran 7y ago
OP here. Thanks for making this connection! I hadn't heard of OFU before, but it sounds interesting. I'm actually working on another post about "optimizing around the optimal" and use multi-armed bandit as an example =). As Jeff Atwood mentions in his post on SOWH, writing about this has been a great way for me to think about the topic. If it turns out that smarter/more experienced people have already thought about/written about this subject at length, I'm not at all disappointed that I'm retracing their steps. In fact, it means that I independently found the right path!
I see this Coursera course on OFU: https://www.coursera.org/lecture/practical-rl/optimism-in-face-of-uncertainty-4Ggk0 https://www.coursera.org/lecture/practical-rl/optimism-in-fa... and what looks like Auer's original UCRL paper http://papers.nips.cc/paper/3052-logarithmic-online-regret-bounds-for-undiscounted-reinforcement-learning.pdf http://papers.nips.cc/paper/3052-logarithmic-online-regret-b... Do you have a preferred source for learning about this? Thanks and Merry Christmas!
- gautamcgoel 7y agoI like these notes on multi-armed bandits (https://arxiv.org/pdf/1904.07272.pdf https://arxiv.org/pdf/1904.07272.pdf), I think they're pretty accessible. Section 1.3.3 discuss the OFU principle. Happy holidays to you as well!
- gwern 7y agoAnother one you might like is the Amazonism 'disagree and commit': https://www.amazon.com/p/feature/z6o9g6sysxur57t https://www.amazon.com/p/feature/z6o9g6sysxur57t I ( https://www.gwern.net/Timing#try-try-again-but-less-less https://www.gwern.net/Timing#try-try-again-but-less-less ) see it as a kind of bandit as well, but specifically, Thompson sampling, because you can interpret groups of people with strong but inconsistent beliefs individually as collectively implementing a Bayesian distribution, and then individuals going off and following what seems to then like the most profitable opportunity is equivalent to Thompson sampling: https://people.csail.mit.edu/pkrafft/papers/krafft-thesis-final.pdf https://people.csail.mit.edu/pkrafft/papers/krafft-thesis-fi...