4 ms·
How does this change when you're heavily punished for ending up with a much lower number (e.g., in finding a partner, biz deal, etc)? Among a set of losing pla
by curo 7y ago
How does this change when you're heavily punished for ending up with a much lower number (e.g., in finding a partner, biz deal, etc)?
Among a set of losing plays, I presume it's 50/50 on whether you turn over every card (and end up with the last one). So a third of the time you end up in the optimal case, another third you end up doing pretty well, and another third you're subject to absolutely random chance. Guess this is an argument for satisficing vs optimizing in most real world applications.
- anonymousiam 7y agoThe conditions of this "game" are that you must decide then and there which option to choose. In the game, you do not have the luxury of viewing all of the data before you make your choice. Real life can be different (as was mentioned in the video re: Kepler's wife).
- repsilat 7y agoYou're right, this algorithm optimizes for "1 point for picking the best option, 0 for anything else", not knowing the distribution. If you do know the distribution, and have some objective function, and samples are "free", you get a different model that you can solve inductively backwards from the case of "the last sample". A different model is that your objective function is still on order, but it gives some points for "closer to best". Maybe you score minus n for picking the nth best sample. In that case the model need not have a known distribution (It's different if you do or don't.) And the strategy us different: if the 99th sample out of 100 is the second best you've seen so far, you should take it. (Which can't possibly be optimal in the original model.)