4 ms·
Has anyone has tried searching for new basic operations, below the level of neural networks? We've been using these methods for years, and I doubt the first maj
by optimalsolver 3y ago
Has anyone has tried searching for new basic operations, below the level of neural networks? We've been using these methods for years, and I doubt the first major breakthrough in ML is the most optimal method possible.
Consider the extreme case of searching over all mathematical and logical operations to see if something really novel can be discovered.
How feasible would this be?
- wodenokoto 3y agoI’m not sure I’d say NN was the first major breakthrough. For many years people considered them too inefficient to compete with SVM, and people genuinely thought kernels was the way to intelligent machines. Today you find researchers claiming that Bayesian nets will outcompete NN. We’ve also seen tremendous success of random forests and other ensemble models. I am sure there are plenty of researchers looking into all sorts of novel ensembles. I think the major breakthrough is ensemble models and with a little bit of cheekiness you can say that NN are ensemble logistic regressions.
- steppi 3y agoI’m not sure if you are aware, but Bayesian neural networks can be actually be well approximated by appropriate ensembles of standard neural networks [0]. The strength of Bayesian nets (including the approximating ensembles) is that they are able to estimate the uncertainty in their own predictions (by generating a probability distribution of possible predictions), at the cost of more computation needed for training and inference. I don’t think it’s ever going to be a matter of Bayesian nets outright outcompeting standard nets though, it’s just another tool in the toolbox if you want a model which “knows it doesn’t know something” and don’t mind the extra compute needed. [0] https://arxiv.org/abs/1810.05546 https://arxiv.org/abs/1810.05546
- Bewelge 3y agoCouldn't that be a way to address the issue of current LLMs hallucinating?
- steppi 3y agoPossibly, but I struggle to reason about Bayesian nets at that scale. I think the level at which a Bayesian net could “know what it doesn’t know” would be regarding uncertainty in what text to generate in a given context, not whether or not the generated text is saying something true. One example could be a prompt in a language not seen in the training data. It could be that some plausible sounding made up thing is likely in a given context. Also, at the end of the day, what you’ll get out of a Bayesian LLM is a sample of several generated texts which would hopefully have more variation than multiple samples from the same standard LLM. I can see it being helpful to see if the different outputs agree or not, but I can’t tell at a glance how well it would work in practice.
- Bewelge 3y agoThanks for the explanation!
- kqr 3y agoThe efficacy of neural networks really boil down to the efficacy of linear operations, which in turn – I suspect – are efficacious because all smooth functions are linear when you look closely enough. That might help with the intuition for why neural networks seem to represent such a fundamentally useful operation. I'm the wrong person to speculate about the future.
- petercooper 3y agoIf you can come up with desired inputs and outputs and have some building blocks in terms of the operations you can perform on them, sure. This happens with https://en.wikipedia.org/wiki/Superoptimization https://en.wikipedia.org/wiki/Superoptimization and various algorithms have been found using such techniques. It also occurs at higher levels of abstraction, for example: https://www.deepmind.com/blog/alphadev-discovers-faster-sorting-algorithms https://www.deepmind.com/blog/alphadev-discovers-faster-sort...
- h0l0cube 3y ago> Has anyone has tried searching for new basic operations, below the level of neural networks? Perceptrons aren’t even a good analogue for biological neural networks. Each dendrite in and of themselves behave something like a multi layer perceptron. Back propagation doesn’t resemble human learning. Biological neurons also have a temporal activation function. Todays most successful ANNs seem focused on raw compute, but they may be missing some of the secret sauce that permits more interesting behaviors. https://braininspired.co/podcast/167/ https://braininspired.co/podcast/167/
- williamcotton 3y agoI can see closer modeling to biological neurons going in two directions, 1.) definite improvements 2.) and engineering inefficiencies that are economically difficult to overcome. A mindless analogy would be fixed wing aircraft with engines vs. wings that flap. And hey, maybe we end up with flapping wings on commercial aircraft at some point, so who knows! Day one of Intro to CE had a slide that is burned into my brain: Engineering = Physics + Economics
- grugagag 3y ago> 1.) definite improvements > 2.) and engineering inefficiencies that are economically difficult to overcome. Improvements can be had in directions we’re not even thinking about and possibly that spark in artificiality that brings it to life in the self autonomous way. I could see that becoming a bad path for us… Engineering inefficiencies will arise when it’s used for the wrong thing and that’s happening a lot when a ton of money is poured in new tools that are used without being properly understood.
- joeythedolphin 3y agoIt seems the value of ML and AI comes from the approximation of the brain, or best-fitting conceptual models that nevertheless today contain high error. As I understand, dendrites vs perceptrons are quite different, but holistically, this is about energent behavior from simple input output networks. Structuralism is about to be put to the test. The better our conceptual models fit the biological behavior, the more "human-like" the behavior. We should expect the field to continue developing practical and better fitting models. The only question, here then, is a "articial ladder of consciousness", where we decide where in this spectrum a being deserves to have rights and not suffer. We may want to start granting rights to our models today to avoid a history of enslaving conscious beings (serious perspective)
- nextos 3y agoKanerva associative architectures have one extra basic operator not found in classical NNs: https://redwood.berkeley.edu/wp-content/uploads/2021/08/Module1_Kanerva_slides.pdf https://redwood.berkeley.edu/wp-content/uploads/2021/08/Modu...
- bordercases 3y agoThere are only sixteen possible maps from Z^2 to Z^2.