5 ms·
I post a link to NEAT here about once a week. The big problem is that NEAT can't leverage a GPU effectively at scale (arbitrary topologies vs bipartite graphs)
by Rhapso 2y ago
I post a link to NEAT here about once a week.
The big problem is that NEAT can't leverage a GPU effectively at scale (arbitrary topologies vs bipartite graphs)
Other than that, it feels like a "road not taken" of machine learning. It handles building complex agent behaviors really well, and pressure to minimal topology results in understandable and reverse interpretable networks.
Its easier to learn and implement than back propagation, the speciation is the most complex but awesome feature. It develops multiple solutions, clustering them to iterate on them separately. An online learning NEAT agent is in practice is an online collection of behaviors, adapting and swapping dominance as their fitness changes.
- Rhapso 2y agoI think the major iteration that will come next is "NEAT in a dream" where a robot or agent online-trains a library of behaviors on an model of the environment constantly updated with the new experiences generated by the dominate behaviors interacting with reality.
- bob1029 2y ago> The big problem is that NEAT can't leverage a GPU effectively at scale I worry how much experimentation we are leaving on the table simply because these paths wouldn't implicate nvidia stockholders in an advantageous manner. Modern CPUs might as well be GPUs when you are running GA experiments from the 80s and 90s on them. We can do millions of generations per hour and populations that span many racks of machines with the technology on hand right now.
- api 2y agoThe more immediate causal reason is that there is no off the shelf high performance hardware to accelerate these paths. GPUs were made for graphics, not machine learning. They just happened to be really good at running certain kinds of models that can be reduced to a ton of matrix math that GPUs do really really fast. Something like Intel Xeon Phi or one of the many-many-core ARM designs I've heard talked about for years would be better for more open ended research in this field. You want loads and loads of simple general purpose cores with local RAM. Put like 4096 32-bit ARM or RISC-V cores with super fast local RAM on a die and transactional or DMA-like access to larger main memory, and make it so these chips can be cross linked to form even larger clusters the way nVidia cards can. This is the kind of hardware you'd want. When I was in college in the early 2000s I played around with genetic algorithms and artificial life simulations on a cluster of IBM Cell Broadband Engine processors at the University of Cincinnati. That CPU was an early hybrid many-core design where you had one PowerPC-based controller and a bunch of simplified specialized cores like a GPU but more designed for general purpose compute. Programming it was hairy but you could get great performance for the time on a lot of things. Computing is unfortunately full of probably better roads not taken for circumstantial reasons. JavaScript was originally supposed to be a functional language-- a real, cleaner one. There were much, much better OSes around than Unix in the 80s and 90s but they were proprietary or hardware-specific. We are stuck with IPv4 for a long time because IPv6 came too late, and if they'd just added say 1-2 more octets to V4 then V6 would not be necessary. I could go on for a while. This has led to a "worse is better" view that I think might just be a rationalization for the tyranny of path-dependent effects and lock-in after deployment.
- RaftPeople 2y ago> When I was in college in the early 2000s I played around with genetic algorithms and artificial life simulations on a cluster of IBM Cell Broadband Engine processors Early/Mid/Late 2000's I was working on a hobby project for artificial life with ga evolved brains. I spent a lot of time investigating options, like the cell processor (I bought a playstation to test with). I also looked at interesting multi-core cpu's that were being introduced. I ended up with a combo of CPU+GPU on networked PC's, which was better than nothing but not ideal.
- bob1029 2y ago> I ended up with a combo of CPU+GPU on networked PC's, which was better than nothing but not ideal. The biggest constraint I've seen with scaling up these simulations is maintaining coherence of population dynamics across all of the processing units. The less often population members are exchanged, the more likely you will wind up with a decoupled population and stuck in a ditch somewhere. Since modern CPUs can go so damn fast, you need to exchange members quite frequently. Memory bandwidth and latency are the real troublemakers. You can spread a simulation across many networked nodes, but then your effective cycle time is potentially millions of times greater than if you keep it all in one socket. I think the newest multi-die CPUs hit the perfect sweet spot for these techniques (~100ns latency domain).
- RaftPeople 2y ago> The biggest constraint I've seen with scaling up these simulations is maintaining coherence of population dynamics across all of the processing units Agreed. For mine I purposefully avoided this problem by making the population+world relatively small (<100 creatures and a relatively finite space they could move in) so each pc handled one instance of a world and after each generation there was a process of sharing data so successful brains were available to each world for continued evolution. My dream was to scale this thing up significantly, but life intervened, maybe someday.
- tomrod 2y agoI feel a bit embarrassed as, despite almost 15 years working in the field, I've not played with NEAT yet nor read up well on it. TIME TO CHANGE THAT :)
- epr 2y ago> The big problem is that NEAT can't leverage a GPU effectively at scale (arbitrary topologies vs bipartite graphs) Is that true? These graphs can be transformed into a regular tensor shape with zero weights on unused connections. If you were worried about too much time/space used by these zero weights, you could introduce parsimony pressure related to the dimensions of transformed tensors rather than on the bipartite graph.
- Y_Y 2y agoOr even use CuSparse, if you don't mind a little bit of extra work over normal cudnn. https://developer.nvidia.com/cusparse https://developer.nvidia.com/cusparse
- PartiallyTyped 2y agoWe can also show that sparse NNs under some conditions are ensembles of discrete subnets, and the authors of the original dropout paper argue that [dropout] effectively creates something akin to a forest of subnets all in "superposition".
- HelloNurse 2y agoAnother plausible strategy to neutralize arbitrary topologies: compile individual solutions or groups of similar solutions into big compute shaders that execute the network and evaluate expensive fitness functions, with parallel execution over multiple test cases (aggregated in postprocessing) and/or over different numerical parameters for the same topology.
- Rhapso 2y agoJust because you can pack the topology into a sparse matrix doesn't make it actually go faster. Sparse matrices often don't see good speedup from GPUs. In addition, each network is unique, each neuron can have an entirely different activation function, and the topology is constantly changing. You will burn a lot on constantly re-packing into matrices that then don't see the same speedups a more wasteful topology pretends to have. On the flip-side out narrative of "speedup" is on bipartite graphs crunch faster in gpus and it might not be the same if the basis is actually utility of behaviors generated by the networks. A cousin thread explores this better.
- nurettin 2y agoI never understood the appeal of NEAT. It's easy to conceive of mutation operators on fully connected layers instead of graphs of neurons and evaluate a lot faster. NEAT also seems to have at least a dozen hyperparameters.
- Rhapso 2y agoIt mostly manages its hyperparameters itself. There is a reward function, one of the hyperparameters we do have to set, for condensing functionality into topology instead of smearing it obscurely across laters. You can just look at a NEAT network analytically and know what is going on there.
- buffalobuffalo 2y agoI've used NEAT a few for a few different things. The main upside of it is that it requires a lot less hyper-parameter tuning than modern reinforcement learning options. But that's really the only advantage. It really only works on a subset of reinforcement learning tasks (online episodic). Also, it is a very inefficient search of the solution space as compared to modern options like PPO. It also only works on problems with fairly low dimensional inputs/outputs. That being said, it's elegant and easy to reason about. And it's a nice intro into reinforcement learning. So definitely worth learning.