3 ms·
I think you are right, from the paper: >During training we used three agent types that differ only in the distribution of opponents they train against, when th
by ogig 7y ago
I think you are right, from the paper:
>During training we used three agent types that differ only in the distribution of opponents they train against, when they are snapshotted to create a new player, and the probability of resetting to the supervised parameters.
So exploiter agents aren't fed a specific strat, instead they discover the weak spots in the same way as the main agent tries to win. The GANs similarity is there.