5 ms·
It's all deep learning, making it the exact opposite of a Cambrian event. Has anyone has tried searching for new basic operations, below the level of neural ne
by optimalsolver 4y ago
It's all deep learning, making it the exact opposite of a Cambrian event.
Has anyone has tried searching for new basic operations, below the level of neural networks? We've been using these methods for years, and I doubt the first major breakthrough in ML is the most optimal method possible.
Consider the extreme case of searching over all mathematical and logical operations to see if something really novel can be discovered.
How feasible would this be?
>Tying it together, a big diversification is likely coming
I'd be interested in knowing what signs, if any, currently point to that.
What I see is a coalescing around the current paradigm. And considering that the hardware side led by NVIDIA is catering to this paradigm, I don't see us breaking out of this local optimum any time soon.
- tarvaina 4y agoRegarding whether this can be likened to Cambrian explosion: We don’t know what caused Cambrian explosion but conceivably it was some evolutionary innovation such as vision. Similarly deep learning is a cause for the explosion of new AI use cases.
- seydor 4y ago> for new basic operations, below the level of neural networks? Why would we? It's proven that neural networks can learn arbitrary functions, and deep networks can apparently learn all the things we don't have functions for. Backpropagation is probably biologically implausible but it's a very good and fast method to train such networks. What could another implementation of intelligence offer that we don't already have? Backpropagation is very old and maybe it's not optimal but it s good enough that it doesnt matter
- version_five 4y agoCan learn arbitrary functions says nothing about efficiency or any practical considerations. A vanilla neural network can't learn a sine function or anything periodic, for inputs outside the range of the training data. Re the question of other methods, Jeffrey Hinton proposed what I think is called a "forward-forward" alternative to back propagation somewhat recently. I don't think it's been proven to be better in any way, but it shows people are searching for new methods.
- seydor 4y ago> says nothing about efficiency or any practical considerations Sure it does, it's practical for pretty much any problem we want to solve, in fact i can't think of something it has not been shown to be practical about. Gradient descent is pretty straightforward and the most efficient generic method we have, even though it's not proven to be the fastest. Forward-forward is proposed as an alternative that is biologically plausible, not as an efficient alternative (it isnt)
- snovv_crash 4y agoIf you have a smaller number of parameters, then quadratic descent (eg. Levenberg Marquardt) is a lot faster for most problems. Gradient descent is only used for NN because we initialize all the parameters to random values, so the assumption of the JtJ approximating the Hessian (required by LM) doesn't hold, since residuals are too high. I suspect there might also be something about the learning rate effectively being too high when using quadratic descent. Even for large problems there are techniques like LBFGS which converge a lot faster than gradient descent, that don't require O(N^2) memory on the number of parameters.
- slabity 4y agoWhy are you responding to only parts of comments and ignoring the context around it? It just makes it look like strawmen arguments. > Can learn arbitrary functions says nothing about efficiency or any practical considerations. This is absolutely true. A turing machine can perform any computation, but they are not really practical for anything outside of an academic exercise. > A vanilla neural network can't learn a sine function or anything periodic, for inputs outside the range of the training data. This is also true. A vanilla neural network isn't that useful outside of very basic tasks with a lot of training data. It's only when you start changing their shapes and using specialized techniques to fit the problem that they become practical. Even a simple task like classifying MNIST digits is done far more efficiently using a basic CNN than a vanilla NN.
- agnosticmantis 4y agoIt’s funny that the masters (Hinton, LeCun, et al) are constantly looking beyond the orthodoxy while the disciples worship the masters’ creations and fight any and all criticism. Wondering where I saw that pattern before…
- mejutoco 4y agoArbitrary functions that are differentiable, unless I am mistaken.
- seydor 4y agobounded iirc, the activation functions need to be differentiable, but in practice they dont need to be strictly
- mejutoco 4y agoI thought gradient descent only worked on differentiable functions. https://en.wikipedia.org/wiki/Gradient_descent https://en.wikipedia.org/wiki/Gradient_descent
- b33j0r 4y agoSometimes I feel like all we’re saying is that we’re surprised that all of our capabilities might just be a statistical model. From everything we do know, what else would it be? The arrogance of consciousness is our blind-spot sometimes. Even the words my brain just typed are combinations of words and arguments from everything you’ve ever agreed or disagreed with. So, in my view. Traceability in AI is possible, but so should traceability in neurology be possible. It might not be as important as it seems to us. And that humility hits people in various ways. My idea to calm everyone down about traceability is to have the AI write a log/wiki in parallel with its activities. Putting toothpaste back in a tube has never worked, so don’t light your own hair on fire, right?
- barking_biscuit 4y ago>Putting toothpaste back in a tube has never worked The confetti has left the cannon.
- lachlan_gray 4y agoIt's not in the networks themselves, but the combinatorial ways that we can string them together. Even if the language models themselves don't become much more powerful, there are a lot of really weird things to discover in terms of how we can pipeline them together to accomplish things. For example things like this: https://yoheinakajima.com/task-driven-autonomous-agent-utilizing-gpt-4-pinecone-and-langchain-for-diverse-applications/ https://yoheinakajima.com/task-driven-autonomous-agent-utili... In a way, you can also look at language models as being the new basic operation. Activations and floating point math are replaced by words and symbolic reasoning.