3 ms·
Agreed that the strategy is an interesting part. Another interesting part will be creating an AI / neural network that can utilize inputs that are closer to hu
by king07828 8y ago
Agreed that the strategy is an interesting part.
Another interesting part will be creating an AI / neural network that can utilize inputs that are closer to human level inputs (e.g., using the frame buffer and audio out as input to the neural network and passing the outputs of the neural network to a keyboard and mouse driver). Just let the network train itself without having a human laboriously determine the topology of the neural network. Such a neural network can then be applied to several different types of games / problems much more quickly than at present where significant human labor is required to generate deeply customized neural networks for each game / problem.
- rhaps0dy 8y ago>Just let the network train itself without having a human laboriously determine the topology of the neural network I hope this little koan illustrates that this sentence is impossible to execute. The human always has to specify something. -- In the days when Sussman was a novice, Minsky once came to him as he sat hacking at the PDP-6. "What are you doing?", asked Minsky. "I am training a randomly wired neural net to play Tic-tac-toe", Sussman replied. "Why is the net wired randomly?", asked Minsky. "I do not want it to have any preconceptions of how to play", Sussman said. Minsky then shut his eyes. "Why do you close your eyes?" Sussman asked his teacher. "So that the room will be empty." At that moment, Sussman was enlightened.
- king07828 8y agoAgreed that without any constraints, it could become a Sisyphean task. The exercise then becomes one of finding the minimal constraints needed to achieve the desired results. Please correct if needed, but looking at the Dota 2 neural network [1], it boils down to generating an input State vector from the Dota 2 bot output interface, running the state Vector through an lstm (of sufficient length) to generate an output State vector, and generating the inputs for the Dota 2 bot input interface from the output State vector. Update this network (1) to have the input State Vector generated from a convolutional network that feeds a fully connected Network and uses the frame buffer as input and (2) to have the final outputs of the neural network be keyboard and mouse commands instead of dota 2 bot input interface commands, then let the network train itself. The number of elements in the state vector, the number of convolutional layers, the number of lstm layers, and the number of layers and elements in each fully connected hidden layer could each also be determined by a recurrent neural network. [1] https://towardsdatascience.com/the-science-behind-openai-five-that-just-produced-one-of-the-greatest-breakthrough-in-the-history-b045bcdc2b69 https://towardsdatascience.com/the-science-behind-openai-fiv... (see the image under "The Architecture") [ random capitalization powered by Google speech dictation ]
- dojomouse 8y agoThe main reasons they don't do this are that it's a fairly known quantity from an ML perspective (going from sequences of images to representational features), so wouldn't be proving that much to be able to do (c.f. the various Atari benchmarks which adequately learned actions to achieve rewards working with pixel inputs)... but at the same time would consume a huge fraction of the computer resource they really want to be targeting at the core timing/tactics/strategy problems... which is where they're really going beyond what's been demonstrated elsewhere with RL. I agree it'll be even cooler when it all justworkstm end to end, but in terms of incremental 'holyshiticantbelievethatworked' this is at least as big a step as it will be when they add in direct visual input.
- king07828 8y agoAgreed. One of the next significant moments could be taking the current Dota 2 algorithm and massaging it to use human style inputs and outputs. Please correct if needed, but the current Dota 2 algorithm boils down to (1) a fully connected network that generates an input state vector from the Dota 2 bot output interface, (2) an LSTM of sufficient length that generates an output state vector from the input state vector, and (3) another fully connected network that generates the Dota 2 bot interface inputs from the output state vector. This could be updated to have (1a) a convolutional network that feeds into a fully connected network, where the input to the convolutional network is the frame buffer (and perhaps the audio output) and the output of the fully connected network is the input state vector, (2) the same or similar LSTM network, and (3a) a fully connected network that outputs keyboard and mouse commands instead of DotA 2 bot interface inputs. It is an open question as to whether current compute power is sufficient for this massage.