3 ms·
You are correct. There is no TD learning in AGZ. The value network is trained to directly predict the game outcome given the current state, and is not trained t
by psb217 9y ago
You are correct. There is no TD learning in AGZ. The value network is trained to directly predict the game outcome given the current state, and is not trained through "bootstrapping" based on the next state's value estimate.