3 ms·
Interesting! Is there some kind of meta-objective for "AI optimal" that could have replaced the 6 or 7 iterations you did with human R&D? For instance, if you h
by ericjang 6y ago
Interesting! Is there some kind of meta-objective for "AI optimal" that could have replaced the 6 or 7 iterations you did with human R&D? For instance, if you had real human playtesters interacting with the prototypes, is there some signal you could extract to measure that it's "good"?
- CJefferson 6y agoThe problem is AIs are very good at optimising what you asked them for, rather than what you meant to ask for, and figuring out what you want is super hard :) As a simple example: * Start by optimising "players can always do something on their turn" -- but that just ends up with everyone always having exactly one thing they can do (no choice). * So then say "give players more things to do each turn" -- but then they end up being able to do everything every turn (the game gives them too much 'money' (still not really a choice) * OK, so we want to force players to make a choice -- so we say "No, give players as much choice as possible, but make sure if they choose an option it blocks off others (in practice, make as many sets of maximal tasks as possible)" -- but then the AI will make sure every turn every player can do (for example) exactly 3 out of 6 things (any 3), and make sure no matter how well or badly they play they still always get to choose 3 from 6, so the game doesn't really progress, or vary. So, what we want is choice, but also variability, and progress, and players to feel like they are effecting the game, but also don't let one player run away too early, but also don't make it just "feel random who wins", etc.
- ericjang 6y agoGot it. I do research in reinforcement learning and I can sympathize with the difficulty here - in my experience, even something as simple as "I want to balance two objectives: 1) the agent should get a high score and 2) the agent should try to make as few decisions as possible" tends to result in the agent doing neither of those things well"