5 ms·
Yes, they exist, and they are called Linear Regression and Decision Tree. Not everything needs to be a neural network. Anyway, residual connections in NNs as w
by donkeyboy 4y ago
Yes, they exist, and they are called Linear Regression and Decision Tree. Not everything needs to be a neural network.
Anyway, residual connections in NNs as well as distillation being only a 1% hit to performance imply our models are way too big.
- heyitsguay 4y agoI think this is one of those issues where it's easy to observe from the sidelines that models "should" be smaller (it'd make my life a whole lot easier), but it's not so clear how to actually create small models that work as well as these larger models, without having the larger models first (as in distillation). If you have any ideas to do better and aren't idly wealthy, I'd suggest pursuing them. Create a model that's within a percentage point or two of GPT3 on big NLP benchmarks, and fame and fortune will be yours. [Edit] this of course only applies for domains like NLP or computer vision where neural networks have proven very hard to beat. If you're working on a problem that doesn't need deep learning to achieve adequate performance, don't use them!
- z3c0 4y agoI've always thought it was abundantly clear how to make smaller models perform as well as large models: keep labeling data and build a human-in-the-loop support process to keep it on track. My perspective is more pessimistic. I think people opt for huge unsupervised models because they believe that tuning a few thousand more input features is easier than labeling copious amounts of data. Plus (in my experience) supervised models often require a more involved understanding of the math, whereas there's so many NN frameworks that ask very little of the users.
- janef0421 4y agoSupervised models would also require a lot more human labour, and the goal of most machine learning projects is to achieve cost-savings by eliminating human labour.
- z3c0 4y agoUp front, yes, but long term, I wholly disagree. A model that performs at 95% or higher will assuredly eliminate human work, no matter how many interns you enlist to label the data.
- heyitsguay 4y agoPeople have tried (and continue to try) that human-in-the-loop data growth. Basically any applied AI company is doing something like that every day, if they're getting their own training data in the course of business. It helps but it won't turn your bag-of-words model into GPT3. Companies like Google have even spent huge amounts of time and money on enormous labeled datasets -- JFT-300M or something like that for computer vision tasks, as you might guess, ~300M labeled images. It creates value, but it creates more value for larger models with higher capacity.
- deleted 4y ago[deleted]
- z3c0 4y agoI "have tried (and continue to try) that human-in-the-loop data growth" to enormous success, bringing logistic regression models to greater than 99% accuracy. And you can chain vectorization strategies to create more input features than simply a bag-of-words, like morphology, shape, etc. We (the software company that I work for) don't need GPT-3, because it is a specialized model geared towards generating human-like text. Most NLP problems are just parsing text for actionable information, and oftentimes, supervised models can be chained to create something far more effective towards your needs than trying to shoehorn a massive general-purpose unsupervised model into a specialized problem.
- mrguyorama 4y agoIt's almost like we have no clue what we are doing with NN and are just tweaking knobs and hoping it works out in the end. And yet people still like to push this idea that we will magically and accidentally build a superintelligence on top of these systems. It's so frustrating how deep into their own koolaid the ML industry is. We don't even know how the brain learns, we don't understand intelligence, there's no valid reason to believe a NN "learns" the same way a human brain learns, and individual human neurons are infinitely more complex and "learning" than even a single layer of a NN.
- heyitsguay 4y agoAs someone in the ML industry, who knows many people in the ML industry, we all know this. It's non-technical fundraisers that spread the hype, and non-technical laypeople that buy into it. Meanwhile, the folks building things and solving problems plug right along, aware of where limitations are and aren't.
- scrumlord 4y ago[dead]
- hooande 4y ago> It's almost like we have no clue what we are doing with NN and are just tweaking knobs and hoping it works out in the end. No, we understand very well how NNs work. Look at PartiallyTyped's comment in this thread. It's a great explanation of the basic concepts behind modern machine learning. You're quite correct that modern neural networks have nothing to do with how the brain learns or with any kind of superintelligence. And people know this. But these technologies have valuable practical applications. They're good at what they were made to do.
- sdenton4 4y agoWe've got lots of great tricks for making audio ml run fast (we need to produce 16k samples per second on a mobile phone CPU, and sound great), but I think they haven't back propagated to the image or language communities.
- heyitsguay 4y agoInteresting, any material you can share?
- PartiallyTyped 4y ago> Anyway, residual connections in NNs as well as distillation being only a 1% hit to performance imply our models are way too big. I disagree with the conclusion. It indicates that our optimisers are just not good enough, likely because gradient descent is just weak. The argument for residual connections is that we can create a nested family of models which enables expressing more models, but also embedding the smaller ones into them. The smaller models may be retrieved if our model learns to produce the the identity function at later layers. The problem though is that that is very difficult, meaning that our optimisers are simply not good enough at constructing identify functions. With the residual layers, we can embed the identity function into the structure of the model, and we now need to learn to map to 0 (since a residual is f(x) = x+g(x)), we need only to learn g(x)=0). As for our optimisers being bad, the argument is that with an overparameterised network, there is always a descent direction, but we land on local minima that are very close to the global one. The descend direction may exist in the batch, but when considering all the batches, we are at a local minimum. We can find many such local minima via certain symmetries. The general problem however is that even with the full dataset, we can only make local improvements in the landscape. Thus, it’s that the better models are embedded within the larger ones, and more parameters enable us to find them because of nested families, symmetries, and because of always having a descent direction.
- visarga 4y ago> It indicates that our optimisers are just not good enough, likely because gradient descent is just weak. No, the networks are ok, what is wrong is the paradigm. If you want rule or code based exploration and learning it is possible. You need to train a model to generate code from text instructions, then fine-tune it with RL on problem solving. The code generated by the model is interpretable and generalises better than running computation in the network itself. Neural nets can also generate problems, tests and evaluations of the test outputs. They can make a data generation loop. As an analogy, AlphaGo generated its own training data by self play and had very strong skills.
- PartiallyTyped 4y agoI did say that the networks are okay. In fact, I am arguing that the networks are even overcompensating for the weakness of optimisers. Neural nets are great even given that they are differentiable and we can propagate gradients through them without affecting the parameters. I don’t think that this reply takes into consideration just how inefficient RL and the likes are. In fact, RL is so inefficient that current SOTA in RL is … causal transformers that perform in-context learning without gradient updates. Depending on the approach one takes with RL, be it policy gradients or value networks, it still relies on gradient descent (and backprop). Policy gradients are just increasing the likelihood of useful actions given the current state. It’s a likelihood model increasing probabilities based on observed random walks. Value networks are even worse because one needs to derive not only the quality of the behaviour but also select an action. Sure enough, alternative methods exist such as model based RL, etc, and for example ChatGPT use RL to train some value functions and learn how to rank options, but all of these rely on gradient descent. Gradient descent, especially stochastic, is just garbage compared to stuff that we have for fixed functions that are not very expensive to evaluate. With stochastic gradient descent, your loss landscape depends on the example or mini batch, so a way to think about it is that the landscape is a linear combination of all the training examples, but at any time you observe only some of them and cope that the gradient doesn’t mess up too bad. But in general gradient descent shows linear convergence rate (cf Nocedal et al Numerical Opt, or Boyd and Vanderberghe’s proof where they bound the improvement of the iterates), and that’s a best case scenario (meaning non stochastic, non partial). Second order methods can get quadratic convergence rate but they are prohibitly expensive for large models, or require hessians (good luck lol). None of these though address limitations imposed by loss functions, eg needing exponentially higher values to increase a prediction optimised by cross entropy (see the logarithm). Nor do they address the bound on the information that we have about the minima. So needing exponentially more steps (assuming each update is fixed in length) while relying on linear convergence is … problematic to say the list