4 ms·
I'm a PhD student working with neuromorphic computing. I like to think about SNNs as RNNs with discretized outputs. The neurons themselves may have some complic
by jegp 3y ago
I'm a PhD student working with neuromorphic computing. I like to think about SNNs as RNNs with discretized outputs. The neurons themselves may have some complicated nonlinear dynamic (currents integrating into the membrane voltage somehow etc.) but they are essentially just stateful transfer functions. The notion of spikes is a crippling simplification, but it's power efficient and you can argue for numerical stability in the limit. So I tend to consider spikes as an annoying engineering constraint in some neuromorphic systems. Brains function perfectly well without them, although in smaller scales (C. elegans).
The true genius of neuromorphics in my view, is that you can build analog components that performs neutron integration for free. Imagine a small circuit that "acts" like the stateful transfer function, with physical counterparts to the state variables (membrane voltage, synaptic current, etc.). In such a circuit you don't need transistors to inefficiently approximate your function. Physics is doing the computation for you!
This gives you a ludicrous advantage over current neural net accelerators. Specifically 3-5 orders of magnitude in energy and time, as demonstrated in the BranScaleS system https://www.humanbrainproject.eu/en/science-development/focus-areas/neuromorphic-computing/ https://www.humanbrainproject.eu/en/science-development/focu...
Unfortunately, that doesn't solve the problem of learning. Just because you can build efficient neuromorphic systems doesn't mean that we know how to train them. Briefly put, the problem is that a physical system has physical constraints. You can't just read the global state in NWN and use gradient descent as we would in deep learning. Rather, we have to somehow use local signals to approximate local behaviour that's helpful on a global scale. That's why they use Hebbian learning in the paper (what fires together, wires together), but it's tricky to get right and I haven't personally seen examples that scale to systems/problems of "interesting" sizes. This is basically the frontier of the field: we need local, but generalizable, learning rules that are stable across time and compose freely into higher-order systems.
Regarding educational material, I'm afraid I haven't seen great entries for learning about SNNs in full generality. I co-author a simulator (https://github.com/norse/norse/ https://github.com/norse/norse/) based on PyTorch with a few notebook tutorials (https://github.com/norse/notebooks https://github.com/norse/notebooks) that may be helpful.
I'm actually working on some open resources/course material for neuromorphic computing. So if you have any wishes/ideas, please do reach out. Like, what would a newcomer be looking for specifically?
- lucubratory 3y agoI don't know anything about SNNs, so I think I'm who you're asking. Something I'm really interested in is if there's any possibility of transferring training from normal NNs to the sort of physically embodied things you're discussing. Like how RWKV trains its weights like it's a transformer but then acts on them like it's an RNN, would it be possible to do training with the sort of large deployed NNs that are all the rage right now, but then somehow instantiate those weights into the hardware you're discussing? I'm guessing it's non-viable as-is because of the discrete nature of SNN function, and rounding up or down probably doesn't work, but I would be interested in anything you have to say on it. Also, I read years ago about a project that was similarly instantiating NNs physically, but it was using optical properties of layered plates to perform the equivalent of weights, do you know anything about that? I don't think it was discrete (can't see why it would be operating on light), but I'd be interested in anything you have to say about that too.
- jegp 3y agoIf we think about spikes as discretized transfer functions, I would say it's totally viable to "hack" them to represent numerical approximations. In fact, a recent paper demonstrates that exact mappings between ANNs and SNNs exist: https://arxiv.org/abs/2212.12522 https://arxiv.org/abs/2212.12522 There are some pitfalls here, though, and I'm biased against these kinds of methods because they don't use the temporal traces of the neuron integration. Regarding RWKV, someone actually trained a "SpikeGPT": https://arxiv.org/abs/2302.13939 https://arxiv.org/abs/2302.13939 That's a neat insight, which will be great for porting these models onto energy-efficient devices. But the learning problem is still the most interesting open question to me. If we crack that, we can scale down GPT-like models by several orders of magnitude since we can "re-learn" subproblems instead of "hardcode" a silly number of permutations, like the present models do. Neuromorphic hardware (brains included) lend themselves incredibly well to learning. We just don't know how to exploit that yet. Regarding the optical layers, are you referring to optical chips like this one https://www.nature.com/articles/s41467-020-20719-7 https://www.nature.com/articles/s41467-020-20719-7 ? That would be an example of using optics to implement your stateful transfer functions (https://en.wikipedia.org/wiki/Optical_neural_network https://en.wikipedia.org/wiki/Optical_neural_network), but there are several of other incredibly promising technologies such as memristors (https://en.wikipedia.org/wiki/Memristor https://en.wikipedia.org/wiki/Memristor), quantum materials (https://arxiv.org/abs/2204.01832 https://arxiv.org/abs/2204.01832) and even biologically based chips (https://en.wikipedia.org/wiki/Wetware_computer https://en.wikipedia.org/wiki/Wetware_computer). My take on this is that these technologies exploit different principles of physics to "compute" in some way. But I like to think that our computational theories and principles are independent of the implementation substrates. There's still a long way to go, but practically speaking, I'm convinced this kind of hardware will have profound consequences for the way that we compute today. We're talking at least 3 orders of magnitude in compute. Imagine ChatGPT running 1000 times as fast. It's ridiculous.