3 ms·
What LeCun is missing is that the IBM chip was designed to be ultra low power (63mW to do pattern recognition on a real time video). One of the novel technique
by oldmanLecun 12y ago
What LeCun is missing is that the IBM chip was designed to be ultra low power (63mW to do pattern recognition on a real time video). One of the novel techniques this chip uses to achieve low power is to use spikes to send information between neurons.
It is not clear how LeCun would propose to send information in his convolutional neural networks in specialized hardware (from what I can tell, he uses FPGAs which are very high power in comparison). In the worst case, that approach sends data between neurons on each time step, which is very inefficient in power. If they do something more clever, I would bet it would start to like sending spikes.
It also seems very short sighted to say that if the hardware is not specifically designed for convolutional neural networks, then it is not the right architecture. It seems like the IBM chip does support convolutional networks, but it might require a few extra time steps to average.
Biology has evolved to use spikes (across many different species). Perhaps evolution didn't get the memo that non-spiking convolutional neural networks is the only architecture worth building. Maybe it will take a few more thousand generations before evolution catches up, but until then spiking neuron architectures seem like a decent gambit ...
- jbarrow 12y agoI tend to agree that evolution reaches a local optimum given enough time, but it seems that this chip is geared towards machine learning rather than biological accuracy. And currently integrate-and-fire spiking neurons don't appear to work better on the data we're interested in. In this light, and although CNNs aren't the only architecture, his criticisms may be a little more reasonable.
- oldmanLecun 12y agoFrom my understanding, the chip is geared towards implementing neural networks with low power consumption, which makes using spikes a reasonable design choice. So one can argue that using spikes is not about biological accuracy, but power efficiency. That is great that people want general purpose machine learning chips, the question is how to do it in low power. My guess is that the right architecture will be a mix of ML primitives mixed in with things like spikes (and perhaps other primitives seen found in biology).
- mjn 12y agoYes, that's also my understanding. Important background is that this is not (or at least not solely) a commercial initiative by IBM to produce a machine-learning chip, though I'm sure they would love to sell some too. It's a DARPA initiative to find a way to greatly reduce the power budget needed for large-scale data-processing. And one of the starting hypotheses of this particular program, SyNAPSE, is that sparseness in time, aka spikiness, is part of why biological organisms seem capable of processing large amounts of video/etc. data with lower power budgets than computers seem to require. Here's an excerpt from their program statement [1]: Current computers are limited by the amount of power required to process large volumes of data. In contrast, biological neural systems, such as the brain, process large volumes of information in complex ways while consuming very little power. Power savings are achieved in neural systems by the sparse utilizations of hardware resources in time and space. Since many real-world problems are power limited and must process large volumes of data, neuromorphic computers have significant promise. That may or may not be a good hypothesis, but it seems interesting to investigate. In any case, LeCun's real beef is with the DARPA program managers: he thinks a different area of ANN research would've been a better allocation of funds, because in his view this is not among the most promising lines of research. Not an uncommon reaction to DARPA choices, and not always wrong either, but DARPA's got the money. [1] http://www.darpa.mil/Our_Work/DSO/Programs/Systems_of_Neuromorphic_Adaptive_Plastic_Scalable_Electronics_%28SYNAPSE%29.aspx http://www.darpa.mil/Our_Work/DSO/Programs/Systems_of_Neurom...
- viscanti 12y ago> Biology has evolved to use spikes (across many different species). Perhaps evolution didn't get the memo that non-spiking convolutional neural networks is the only architecture worth building. I think this is representative of a lot of AI now. This chip doesn't obviously improve the state of the art on an arbitrary (but standard) benchmark, so LeCun dismisses it. That type of attitude strikes me as over fitting for an arbitrary benchmark (a local optimum) and missing the bigger picture. This chip (and line of research) could help identify how the brain works (which may very well also unlock strong AI).
- bhc 12y agoDid you make a throwaway account just to mock Lecun without people finding out who you are?
- Udo 12y ago> Biology has evolved to use spikes (across many different species). Those different species didn't independently arrive at their own neuronal implementation, we're all using the same basic pattern (with heavy modifications). More importantly though, in biology, most cells use temporal encoding - one spike isn't meant to be a binary event. Instead, the frequency of pulses over time is used to encode intensity. The receiving end can then continuously integrate the signal over time. I'm not sure the IBM chip uses spikes in the same way. > Perhaps evolution didn't get the memo that non-spiking convolutional neural networks is the only architecture worth building. Temporal encoding in nature is used for the obvious reason of hardware efficiency. If every axon had 8+ binary signal lines it would become unreasonably complex and large, and prone to failure. This somewhat mirrors the design decision in the TrueNorth chip, again for obvious reasons. However, the two diverge algorithmically. In nature, signal processing is used to transport relative values over the network. The chip on the other hand really seems to use boolean encoding. You're right about one point though: it's fairly certain that a large number of configurations and algorithms do produce "working" neural networks. Nobody said that imitating our own implementation details will lead to better results than anything else we might want to try. And in fact, that's the reason why so many computational implementations of neuronal networks do not use temporal encoding, because encoding values in floating points or even integers comes more naturally to computers than it does to wetware. But it is an abstraction that takes up comparatively many resources.