5 ms·
An alternative derivation using dynamic programming for multi-round bets: https://adityam.github.io/stochastic-control/mdp/optimal-gambling/ https://adityam.git
by adiM 5y ago
An alternative derivation using dynamic programming for multi-round bets: https://adityam.github.io/stochastic-control/mdp/optimal-gambling/ https://adityam.github.io/stochastic-control/mdp/optimal-gam...
- OscarCunningham 5y agoThe dynamic programming and multi-round bets are a distraction there. Since it uses a logarithmic utility function the Kelly criterion is optimal for even a single bet.
- bumbledraven 5y agoNo, the two betting strategies diverge. The Kelly criterion assumes an unlimited steam of future bets, while the dynamic programming approach assumes a fixed number of remaining bets.
- OscarCunningham 5y agoThe logarithmic utility function makes the Kelly criterion optimal even when there is only one remaining bet. For example say you have $x, the probability of winning is p, and you bet $y. Then you want to maximise plog(x+y) + (1-p)log(x-y). Setting the derivative with respect to y equal to 0 yields p/(x+y) - (1-p)/(x-y) = 0. This rearranges to give y = (2p - 1)x, which is precisely the Kelly criterion.
- adiM 5y agoThis is precise the argument at the penultimate time-step in the dynamic programming solution of the multi-round case. The other interesting aspect is that the expected returns are logarithmic, i.e, with y = (2p -1) x p log(x+y) + (1-p) log(x-y) = log(x) + C where C is the Shannon capacity of the binary symmetric channel with cross-over probability p. By the same argument, the expected wealth after T rounds will be log(x) + T C So, in addition to the optimal strategy, we have also derived the rate of growth of wealth. This is also in tune with the motivation of Kelly's paper where he was showing a relationship between Shannon capacity and optimal gambling (without using a dynamic programming argument)
- adiM 5y agoWe can consider the limit T -> infinity to recover the setting with infinite number of bets (but need to look at the rate of growth, otherwise multiple strategies can give infinite returns).
- bumbledraven 5y agoCorrection: I should have said that the Kelly criterion and DP-based approaches diverge when there is a maximum amount of wealth that can be attained (e.g., $250 in TFA), not in scenarios where there are a fixed number of betting rounds remaining.
- adiM 5y agoIn the result for multi-round bets, if we take the betting horizon T=1, we recover the result for a single bit. But the other way round is not obvious.