4 ms·
I've tried out and written about[1] KANs on some small-scale modeling, comparing them to vanilla neural networks, as previously discussed here: https://news.yco
by Ameo 2y ago
I've tried out and written about[1] KANs on some small-scale modeling, comparing them to vanilla neural networks, as previously discussed here: https://news.ycombinator.com/item?id=40855028 https://news.ycombinator.com/item?id=40855028.
My main finding was that KANs are very tricky to train compared to NNs. It's usually possible to get per-parameter loss roughly on par with NNs, but it requires a lot of hyperparameter tuning and extra tricks in the KAN architecture. In comparison, vanilla NNs were much easier to train and worked well under a much broader set of conditions.
Some people commented that we've invested an incredible amount of effort into getting really good at training NNs efficiently, and many of the things in ML libraries (optimizers like Adam, for example) are designed and optimized specifically for NNs. For that reason, it's not really a good apples-to-apples comparison.
I think there's definitely potential in KANs, but they aren't a magic bullet. I'm also a bit dubious about interpretability claims; the splines that are usually used for KANs don't really offer much more insight to me than just analyzing the output of a neuron in a lower layer of a NN.
[1] https://cprimozic.net/blog/trying-out-kans/ https://cprimozic.net/blog/trying-out-kans/
- smus 2y agoNot just the optimizers, but the initialization schemes for neural networks have been explicitly tuned for stable training of neural nets with traditional activation functions. I'm not sure as much work has gone into intialization for KANs I 100% agree with the idea that these won't be any more interpretable and I've never understood the argument that they would be. Sure, if the NN was a single neuron I can see it, but as soon as you start composing these things you lose all interpretability imo
- Lerc 2y agoThis is sort of my view as well, most of the hype and the criticisms of KANs seem to be fairly unfounded. I do think they have a lot of potential, but what has been published so far does not represent a panacea. Perhaps they will have an impact like transformers, perhaps they will only serve in a little niche. You can't really tell immediately how refinements will alter the usability. Finding out what those refinements are and how they change things is what research is all about. I have been quite enjoying following https://github.com/mintisan/awesome-kan https://github.com/mintisan/awesome-kan progress and seeing the variety of things being tried. I have a few ideas of my own I might try at sometime. Between KANs and fixed activation function networks there is an entire continuum of activation function tuning available for research. Buckets of simple parameter activation functions something like xsigmoid(mx) ( ReLU when m is large, GeLU at m=1.7, SiLU at m=1). This adds a small number of parameters for presumably some game Single activation functions as above per neuron. Multi parameterizable activation functions, in batches, or per neuron. Many parameter function approximators, in batches, or per neuron. Full KANs without weights. I can see some significant acclaim being awarded to the person who can calculate a unified formula for determining where additional parameters should go for the largest impact.
- sigmoid10 2y agoMy big issue with KANs is that MLPs can trivially be made functionally identical to them up to an arbitrarily small error. Just take a group of neurons/layers and thanks to the UAT you can get them to model any reasonably well behaved activation function. Now redefine that group as a KAN node and you have something that works exactly the same way. In that sense it is actually strange that KANs with the same number of parameters don't outperform MLPs. This could be seen as a hint that activation functions are not really what matters in the end. This is also something that the biology of real neural networks seems to suggest from experiments with rats. Although there is far too little conclusive research in that area. I'm about 50:50 on whether this question will be solved by biologists or computer scientists.
- Lerc 2y agoWell obviously a function approximator that uses a sum of functions can be produced by an assembly of function approximators. I don't think anyone is going to argue otherwise. The measure of parameter efficiency is the issue at hand. Given the relative newness of this approach even hitting the bar of "sometimes better, under specific circumstances" is impressive. I am less likely to believe in biological research revealing new techniques. I feel like the inverse is more likely, that discovering new techniques will allow biologists to identify mechanisms where it was already in use.
- sigmoid10 2y ago"Relative newess" is a very shaky foundation for any argument. After all, in practice you could use all of the gradient descent methods we have today, you only need to hardcode a few derivatives. On the other hand biology suggests that activation functions are kind of irrelevant for complex networks beyond a certain size. So even in the case KANs slightly better, MLPs would win due to their simplicity in the implementation.
- vinnyvichy 2y agoHaha you are wrong, I have a PhD, and a number of MSc, from the U.S.., institutions that have a number of Nobel laureates.. (but I don't have student debt!) Otherwise -- in a different world -- you and I could be brothers.. (If you judge me to be delusional.. however, erm.. frater quoque! The delusionality helps, sadly)
- wanderingmind 2y agoReally detailed work. Thank you. For those looking to jump straight to code, here is the link to codebase discussed in the blog. https://github.com/Ameobea/kan https://github.com/Ameobea/kan
- alexnewman 2y agoI’m very happy to hear someone else say the quiet part out loud . Everyone claims nn aren’t interpretable, but that’s never been my experience . Quiet the contrary
- DonHopkins 2y agoI call bullshit. Occam's razor and common sense says you simply have no idea what you're talking about. You certainly don't have any proof or explanation. And your other posts about how you support accused pedophile and crypto scammer and Jeffrey Epstein associate Brock Pierce prove that you're a liar, so I'm quite skeptical about all the other outrageously sketchy unsupported claims you make. Care to explain what your "experience" proves, and provide some evidence?