8 ms·
I've recently become interested in permutation invariant neural networks. There has been very little work in this area - just PointNet and a few derivatives. A
by calebh 8y ago
I've recently become interested in permutation invariant neural networks. There has been very little work in this area - just PointNet and a few derivatives.
Anyway, I think that neural networks are now entering the trough of disillusionment as people begin to discover the limitations. Maybe in the future, somebody will come up with a new machine learning architecture that has better generalization. I'm not expecting gradient descent to give us general AI.
- salawat 8y agoThe main cause of the generic brittleness of Neural Networks is probably in the way they are utilized. Biological neural nets never really stop learning. They slow down, even "forget" in order to restructure, but they change constantly. A static neural net in basically a snapshot of it's environment (training data). Very interesting consequences for the ML field if my hunch has anything remotely resembling a kernel of truth to it.
- sova 8y agoContinue learning while also continuing to "forget" ... that's very interesting! Space required to save data remains constant, and their ability adapts to environment and conditions. So... seeing neural nets as living structures instead of as snapshot sieves/filters... very strong approach. I wonder if for solid AI we need to make ai-biology first. To be concrete about what I'm saying now, consider that we are not in conscious control of the generation of our skin cells or the processes of our kidneys, but they do organically persist on their own. Maybe we need to get ai/biology so tight-knit first, so that we can have a proper vehicle for an AI-mind to journey within, while also not necessarily rewriting the code for its heartbeat when it just wants to remember names of common phenomena.
- xapata 8y agoThat's a factor, but it doesn't explain the phenomenon that human-imperceptable transformations of an image can dramatically shift a NN's outputs.
- salawat 8y agoAgain, look to biology. That thing you are modeling. Humanly "imperceptible" is a very loaded term. Human perception has billions upon billions of networks worth of filtering going before we even boil down our environment to the "interesting" stuff. Furthermore, if you take a snapshot of that network after training, you're fit to the training data. The network has lost it's plasticity. Take a potato, put it on the ground, train the network on other potato shots. Now show it a potato shaped asteroid. Now show it a French fry. What is the potato-ness that this potato+detector is ACTUALLY homing in on? Keep in mind, this structure is trained on digital encodings of maps of light and color. The function may not be a perfect semantic detector of potato-ness. It just knows what patterns of bits MIGHT be potatoes. And when you are working on bit level encodings, one bit translates to a lot of change, even if it is imperceptible to a human looking at a rendering on a screen. Heck, there is no guarantee that the function it's emulating is well defined outside the training data set. Neural networks are GOING to be fickle. You're trying to coerce "reliable, repeatable. generifiable results" out of a simulation of the same stuff that drives five year olds and emotional people. Consider yourself lucky the program hasn't opened the CD tray and demanded you insert crayons.
- xapata 8y ago> You're trying to coerce "reliable, repeatable. generifiable results" out of a simulation of the same stuff that drives five year olds and emotional people. The "neural network" machine learning technique is not a simulation of biology. It's a nice marketing phrase. The technique is just math, maybe "inspired" by someone thinking about neurons. Don't be misled by branding. For a good explanation of why some of these fanciful science terms come about, read Bellman's explanation of why he called his research "dynamic programming".
- nightcracker 8y agoI like "differentiable programming" instead of "neural networks". The crux is that you are programming in such a manner that you can differentiate your function w.r.t. parameters which are then tuned.
- sova 8y agoWell said. Do you think some sort of other patterning mechanism, like that used to identify songs (thinking the Shazam algorithm where there is more along the lines of making a "fingerprint" and comparing to the print) will help take us to the next level? Neural Network architecture (sideways skyscrapers that get more focused as the network plays out) does not incorporate generalization much and this is sort of the point: a specific input data shall generate the resulting target, and any small variance would decidedly result in something different at the other end of the NN. Unless there were some way to incorporate a broad pattern on the data as it comes in... For example, run the NN, the NN shows results and also produces a fingerprint hash, now when you put slightly mal-aligned data into the NN, also provide the previous fingerprint result, and now (maybe) we can derive a device that will result in the same product and similar products within a range of the input data, to satisfy the case that the fingerprint stays true even when the data is slightly different. Just some ideas on the topic, if we could solve pattern generalization where the viewing window did not have to line up perfectly with the perceived pattern, that would be very powerful; we would have the ability to measure unaligned states, but an algorithm that approaches this power also must exploit some natural phenomenon, such as quantum superposition, if we are to complete our iterations in reasonable time. Without some sort of preprocessing step, such as sorting (that can ensure algorithms run quickly) it seems difficult to try and make a real neural net that is significantly more than a very specific compressed archive. What's really interesting is that as a stochastic process we could both generate two separate, functional, result-bearing neural networks (weights and layers) that achieved same or similar results while being completely different at their basis (in number of layers and the actual activation weights). And right now, there's no way to determine [elegantly] how similar our neural nets are. So, perhaps in some ways, the ability to compare neural network guts meaningfully will lead us to the ability to encode more elegant view ranges on data sets.
- candiodari 8y ago> Neural Network architecture (sideways skyscrapers that get more focused as the network plays out) does not incorporate generalization much ... Have you ever seen kids learn the tables of multiplication ? Or addition (but people don't seem to remember from their past that they learned addition from tables). Doesn't that give you pause when you're saying humans can generalize ? Because clearly, this is not how we teach our children. Humans don't learn that if 2+2=4, 3+2 must be 5. No. Humans learn that 2+2 is 4, then they learn that 2+3 is 5, then they learn ... and so on and so forth. Even when you see online players investigate a game. Whether it's chess, zelda, or starcraft. You keep seeing the same thing. Humans don't learn by generalizing. They learn how to act in every situation ... one situation at a time. So it's not just 6 year old kids doing this. The game advances because there's a few specific individuals that spend entire days making nonsensical moves, and on rare occasions they find a new move, see what happens, and tell the world (in chess we're talking something like once a year, in mario it seems possible to do it once every few months. It's scary how much tenacity is required for such a process). And for the vast majority of humans, it never even gets to that point. They just don't have the tenacity. So if you think humans are AGI, which you seem to do, then clearly such a state can be achieved with disappointing generalization capabilities.
- rdlecler1 8y agoThis is probably the fifth or sixth trough. The problem is we spend too much time trying to brute force engineer our way forward when we should spend more time reverse engineering the salient properties that give rise to biological intelligence. Someone will invariably say that planes don’t use bird wings, but both birds and planes relying on the laws of aerodynamics and we need some kind of similar theoretical framework for computational intelligence—reverse engineering is the shortest path forward. There are a few moves in this direction but without a better framework for computational intelligence, which we can get to faster through reverse engineering, we’re going to keep hitting these walls. This road includes methods for artificial neurogenesis, genetic encoding of the developmental program at the level of neuron and axon driven by artificial gene regulatory networks (AGRN) evolutionary methods to evolve populations, and adversarial ecosystems for open ended evolution. Unfortunately none of these will give you a quick win....
- bitL 8y ago> but both birds and planes relying on the laws of aerodynamics Complexity of biological systems is so mind-blowing that we likely don't even have computational capacity to simulate one of those processes in veracious detail. In addition, we can't even comprehend primitive, 1000-layer deep neural networks, not mentioning significantly more complex neurons (that are performing some local, protein-based computations alongside electric spikes we measure). We should be happy that Deep Learning surprisingly somewhat works, obliterating "classical" ML and make use of it to the full extent of its capabilities. Romantic idea that we can construct AGI with it should be retired.
- rdlecler1 8y agoThis is lazy thinking. You could say the same thing about planetary motion but the ellipse equation does a pretty good job giving us a null hypothesis to build up from. If you’re a materialist then you’re going to believe that all of the minute biological details are relevant. If you’re a functionalist you can work at a more abstract level. From the 12 years of research I did in this space (both CS and Bio) I can tell you for certain that we can learn a lot more from a functionalist approach. And the reason we “can’t understand” 1000 layer deep neural networks is again because we don’t have a good theory of these types of computational systems. For instance I can tell you that most of the interactions w_ijk are completely spurious and if you cull the network you actually end up with a circuit that you can understand in the same way any EE undergrad can identify a 3 bit adder from the circuit topology.
- 317070 8y agoThere are more permutation invariant networks. Look for Bruno or for Neural Statistician. Both of these are RNN's with the exchangeability property. This means inputs and output can be permuted with each other. If you need something with less strong constraints (but still invariant under permutation of the input), there is also Deep Sets.
- calebh 8y agoAh, thanks for the tip. When I'm reading research papers I often find clusters of related papers with the same keywords. Often finding the separate clusters is a difficult challenge, even if you use Google Scholar and follow all the citations.