6 ms·
Neural networks can approximate any function, but that doesn’t mean they do so efficiently. Depending on the function, they can require incredible amounts of ne
by steppi 4y ago
Neural networks can approximate any function, but that doesn’t mean they do so efficiently. Depending on the function, they can require incredible amounts of neurons and training. At their worst, they devolve into a lookup table. It’s not hard to find these examples either. Just try training a neural network to compute sin(x)!
This is possible! One of the cool things about neural networks is that you can try to encode prior understanding into either the structure of the network or choice of activation function. See the paper Neural Networks Fail to Learn Periodic Functions and How to Fix It by Ziwin, Hartweg, and Uweda. Where they propose the activation function f(x) = x + sin(x)^2 that can encode an understanding that the underlying function should be periodic.
[1] https://proceedings.neurips.cc/paper/2020/file/1160453108d3e537255e9f7b931f4e90-Paper.pdf https://proceedings.neurips.cc/paper/2020/file/1160453108d3e...
- thomasahle 4y agoWhile `x + sin(x)^2` may be monotonic itself, it only takes a simple linear combination of two neurons like `x+sin(x)^2 - (x/2 + sin(x/2)^2)` before you have a completely crazy loss landscape. I have a feeling this is why such activation functions haven't become standard.
- whatshisface 4y agoIf we're allowed to do that, why can't my activation function be sin(x)?
- deleted 4y ago[deleted]
- blackcat201 4y agoYou can use sin as activation function, but that would require careful initialization to avoid gradient explosion as you would ended up with a lot of points where gradient is simply zero. You can refer to Implicit Neural Representations with Periodic Activation Functions for more details.
- zarzavat 4y agoIt works perfectly if you don’t have any parameters.
- yobbo 4y agoIt can, but sin(x) has infinite number of extremes, and the gradients will vanish at those points. Activations will get stuck at 1 and -1 (x=π/2, 3π/2, ...). They set x+(1/a)*sin²(x) to be monotonic, which fixes this. Or you need to optimize without using gradients.
- PartiallyTyped 4y agoYou can actually do that. https://www.vincentsitzmann.com/siren/ https://www.vincentsitzmann.com/siren/
- sicp-enjoyer 4y agoI'm surprised to see almost no discussion of fourier series in that paper, considering fourier series is all about representing signals as linear combinations of sinusoidal functions.
- PartiallyTyped 4y agoYou may be interested in [1] where they go to a great extend to show that the convolution operation that we consider in DL is the dual of fourier series [2]. [1] https://geometricdeeplearning.com https://geometricdeeplearning.com [2] https://arxiv.org/pdf/2104.13478.pdf https://arxiv.org/pdf/2104.13478.pdf page 27 (23 if you count book pages).
- amelius 4y agoIs convolution in DL not implemented with the FFT as the underlying workhorse?
- PartiallyTyped 4y agoProbably no since FFT is slower and less parallelizable than products.
- anoy8888 4y agoI like your article and went to your home page to find more good articles and I like what I saw. Thank you for sharing. The only thing is that I am reading from my phone and the site is not very mobile friendly
- SemanticStrengh 4y agoAnd then your NN can't represent anything else than periodic functions.. If we had to build separate programs for each product requirements variations.. programming would not be viable. More generally neural networks can't even imitate a dumb calculator without throwing absurd errors despite the rules of calculus being trivial and well defined. And matching a calculator is a task order of magnitudes easier than the semantic Causal reasoning abilities of human NLU that is involved in argumentation, inferences and understanding. But people are pathetically fooled by the fallacy of it being a uNiVeRsAl aPpRoXiMaTor.
- telchar 4y agoAnd yet neural networks can solve symbolic integral and derivative problems and differential equations better than other computer algebra programs. Sure one network might fail to compute sin(x) numerically, but another could easily tell you its derivative is cos(x). Turns out they are pretty flexible. Do they need to do everything?
- ABeeSea 4y agoSource? I would be very surprised if there was a neural symbolic PDE solver better than what’s in wolfram mathematica.
- SemanticStrengh 4y agoRelated https://arxiv.org/abs/1806.07366 https://arxiv.org/abs/1806.07366
- mhh__ 4y agoAt the cost of how many parameters?
- viraptor 4y ago> And then your NN can't represent anything else than periodic functions Significant parts of x+sin are close to linear. You don't need to use this activation as the result layer either. Why would we lose anything? > If we had to build separate programs for each product requirements variations We pretty much do? We use both extremely generic frameworks both in technical sense (.net) and organisational (sap). But we also have software written to specific requirements where needed (there's millions of very specific ways to invoice someone, companies get invoicing platforms written just for them from scratch). There's space for both approaches.
- zarzavat 4y agoIt’s hard to walk if you can’t feel your legs. The problem is improperly evaluating the neural network as a function of time, instead of evaluating the network as a function of previous state. When we humans approximate functions (let’s say you’re drawing it on a piece of paper, or waving your arm around) we do not simply look at a clock and feed forward that information directly into our motor neurons. Rather, we have sensory neurons that feed in the current state of the function we are approximating as an input, then approximating the next output of a periodic function becomes trivial. It’s very easy to train a neural network to look at a piece of a sine wave and predict what the next value should be - the fact that the function is periodic actually helps you.
- omrjml 4y agoPeriodic function like sin(x) are not a dynamical system so its previous state does not determine the current state. So it should be approximated in that way.
- doubleunplussed 4y agoPeriodic functions like sin(x) are the solutions to differential equations like dy/dx = -y that describe for example, oscillations of springs, to name but one of an extremely large number of dynamical systems that behave this way.
- omrjml 4y agoOf course sine can appear in the solution for dynamical systems but the function itself is not dynamical. When evaluating sin(x) you do not need to know about the previous state.
- zmgsabst 4y agoThat is true, the problem is with the conclusion you drew from that fact: > So it should not be approximated in that way. We can’t conclude a dynamic approximation is a bad approach based purely on the fact the underlying function isn’t dynamic. The function might nevertheless be easily approximated via dynamics — as in the case of predicting sine from seeing the recent history.
- lukaszwojtow 4y agoShameless plug as I'm the author: Primeclue uses math functions to express models and one of its function is sine. Like here: https://github.com/lukaszwojtow/primeclue/blob/68e3b4c8e9f1ac8f5e4adc16934491135cb66a83/backend/primeclue/src/exec/functions.rs#L107 https://github.com/lukaszwojtow/primeclue/blob/68e3b4c8e9f1a...
- em500 4y agoI have the creeping feeling that soon the field is going to reinvent all of Fourier analysis from scratch.
- orbifold 4y agoIt will be called „Neural Harmonic Analysis“, it will cite one paper by a Russian and otherwise ignore any prior work which didn’t include the word neural network.
- srean 4y agoIts sad that this is not far from the truth
- guidopallemans 4y agowavenet and transformer models already come dangerously close
- PheonixPharts 4y ago> Just try training a neural network to compute sin(x)! This is also the classic demo case for LSTMs, I have a notebook open right now that as a LSTM learning the sine function quite well with 32 dimensional state vector. However the authors point still stands. Neural Networks are not great at computing arbitrary functions where the output is an unbounded real number. The sine function is still limited to range that is bounded to [0,1]. A better example of where NNs really fail is when trying to learn tricky to implement the normal quantile function (inverse CDF). This would be an excellent place for function approximation because a.) generating training data is easy (just run random numbers into the CDF and reverse these arguments into a NN) b.) manually writing the quantile function from scratch is a pain since it involves the inverse error function which is very annoying to implement from scratch. You can learn the standard quantile function for mean=0, sd=1, however if you try to generalize this to taking not only the desired quantile but an arbitrary mean and standard deviation you will not learn anything useful. It's a bit of a shame that neural networks are weak in this area because it would be incredible to have a good tool to approximate inverse functions in general. The fact that we almost never see neural networks being used as a tool for this type of work is evidence of this limitation. In general if you're problem can't be modeled where the output is some vector of probabilities it's not a great fit for NNs.
- steppi 4y agoIt's a bit of a shame that neural networks are weak in this area because it would be incredible to have a good tool to approximate inverse functions in general. The fact that we almost never see neural networks being used as a tool for this type of work is evidence of this limitation. It's funny you say. I haven't actually used NNs much in my research but my background is in math and in my spare time I'm a maintainer for SciPy working on special and statistical functions, often the exact kind of stuff you mentioned. I might take up your challenge and write a blog post about it or something if I get any success. Anyway, I wasn’t trying to invalidate the author’s general point, just point out a fun fact about his example.
- PheonixPharts 4y ago> I might take up your challenge and write a blog post about it or something if I get any success. Success or failure, I would really enjoy seeing that write up! I would be even more excited to be proven wrong. The promise of "universal function approximator" is very temping. My personal dream would be to have it so one could essentially run scipy in reverse and learn the entire library with a NN. Even in the case I gave, the idea that you could learn an arbitrary quantile function means you could also arbitrarily learn a sampler for any distribution, since all you have to do compose a uniform sampler with whatever quantile function you learned. Of course for this example I'm using "solved" using a similar approach with variational inference (pyro has a write up on it: https://pyro.ai/examples/svi_part_i.html https://pyro.ai/examples/svi_part_i.html, you might find David Blei's "Variational Inference: A Review for Statisticians" useful as well https://arxiv.org/abs/1601.00670 https://arxiv.org/abs/1601.00670)