7 ms·
Why are LLMs general learners?
- Paul-Craft 3y agoThis seems like just another way of saying that when you train an LLM on a text, its weights incorporate the tokens in that text, which is nothing really profound. I think the real magic here comes from the fact that LLMs are a specialized sort of neural network, and that neural networks are universal approximators [0]. In other words, LLMs are general learners because they are neural networks. This is also not particularly profound, except that there are mathematical proofs of the universal approximation theorem that give us insight into why it must be so. --- [0]: https://en.wikipedia.org/wiki/Universal_approximation_theorem https://en.wikipedia.org/wiki/Universal_approximation_theore...
- im3w1l 3y agoA lot of universal approximators are piss poor at general learning. It's taken a lot of hard work and clever people to get LLM's to where they are. It's not as simple as neural network and done.
- Retric 3y agoCurrent LLM’s are also piss poor general learners, they are however really good at learning specific things which people value highly.
- im3w1l 3y agoSome 15 years ago, textbooks taught that multi level perceptrons (fully connected feed forward network) with one hidden layer were sufficient because they were universal approximators. That thought kinda held back the field for a long time. Going against that dogma was so revolutionary that new paradigm was given its own name: deep learning. Just because you can find some gotcha counterexample LLM's struggle with doesn't invalidate that we've come a very long way.
- mehh 3y agoNah keep hearing this, was doing multilayer in 90s, the problem was my machine didn’t even have a floating point unit, had to hand roll my own fixed point math and cpu was about 100mhz
- ChatGTP 3y agoWhat's the destination, we've come along way, and where do you think we're going?
- Retric 3y agoI think that was largely a misunderstanding. 20+ years ago I took an AI class that mentioned using multiple levels was useful for training neural networks. It also mentioned a 2 layer network was only a universal approximator given arbitrarily large numbers of nodes which again seems to be forgotten about. Though the teacher worked in industry for a while which may have been relevant as we didn’t focus that much on theory. PS: Deep learning was also more about improving computational power than some major theoretical advancement.
- Paul-Craft 3y agoCorrection: hard work, clever people, and massive increases in computational power. I'm sure all three matter quite a lot here. I'm not saying that if your goal is to come up with a usable general learning algorithm that it is just "as simple as neural network and done." What I'm saying is the converse: that the general learning capabilities of LLMs are most likely explained by the fact that, well, they are general learners, via the universal approximation theorem. Your other comment, I think, suggests why we're just now starting to see more general learning capabilities out of neural networks, when the theory says that a single hidden layer is enough: with a single hidden layer, you really need to get all the weights pretty close to "right" to see general learning/universal approximator behavior. When you have more than one hidden layer, then some of your weights can be wrong, as long as the errors are corrected in later layers. Now, I'm not an AI researcher or even anyone who works anywhere near this area, but I did take a course or two in grad school, and this seems at least intuitively plausible to me. If there are researchers in the field reading this, I'd definitely like to hear their takes, because I'm totally open to being completely wrong here. I'd rather be one of the lucky 10,000 than just have this half-baked idea that seems right. :-)
- butyEah 3y agoHardware matters most. No matter how clever there’s no storing such large parameter sets on an Intel 286 with 4MB RAM. No matter how clever the programmer there’s no encoding GPT4 with that. It was the hardware constraints that required programmers to be clever to begin with. These days it’s much more “copy paste the math directly because our data set is so robust and our hardware and networks so performant clever low level hacks don’t matter.” Especially at big tech where they’ve used their own AI to guide them; the ability to just ask an ML system to simplify math has existed for a few years now, we’ve all seen how clever outputs were set aside for safe linear hacking. Truly clever work is occurring in more traditional sciences like chemistry and biology these days.
- tarvaina 3y agoThe ingredients you need for training a useful machine learning model are expressivity, learnability, and generalization. Many methods are universal approximators but that only takes care of the first ingredient. Arguably the reason neural networks are so successful is that they can offer a good balance between the three. Before transformers we built different neural network architectures for each domain. These architectures offered better inductive biases for their respective domains and thus traded off some of the expressivity for better learnability and generalization. Nowadays the best architectures seem to be merging towards transformers. They appear to offer more generally useful inductive biases and thus a better trade-off between the three ingredients than the earlier architectures.
- dgreensp 3y agoLLMs are not particularly good at arithmetic, counting syllables, or recognizing haikus, though, because (contrary to the thesis of the article) they don’t magically acquire whatever ability would “simplify” predicting the next token. I don’t feel like the points made here align with any insight about the workings of LLMs. The fact that, as a human, I “wouldn’t know where to start” when asked to add two numbers without doing any addition doesn’t apply to computers (running predictive models). They would start with statistics over lots of similar examples in the training data. It’s still remarkable LLMs do so well on these problems, while at the same time doing somewhat poorly because they can’t do arithmetic!
- iliane5 3y ago> LLMs are not particularly good at arithmetic, counting syllables, or recognizing haikus I suspect most of this is due to tokenization making it difficult to generalize these concepts. There are some weird edge cases though, for example GPT-4 will almost always be able to add two 40 digits number but it is also almost always wrong when adding a 40 digit and 35 digit number.
- rcme 3y agoIt doesn't have anything to do with tokenization. You can define binary addition using symbols, e.g. a and b, and provide properly tokenized strings to GPT-4. GPT-4 appears to solve the arithmetic puzzles for a few bits, but quickly falls apart on larger examples.
- iliane5 3y agoWhat I was saying is that because you need to go out of your way to make sure it's tokenized properly, I wouldn't be surprised if there are enough non properly tokenized examples in the dataset. If that was the case, it would make it difficult to generalize these concepts.
- doctor_eval 3y agoCould it also be that syllables are intrinsically mechanical? They are strongly related to how our mouths work. While it may be possible to extract syllables from written text - following the consonants and vowels - I'm not sure that many humans could easily count syllables without using their mouths.
- mxkopy 3y ago> Yet, they demonstrate a crucial point: a deeper understanding of reality simplifies next-token prediction tasks. I'm not sure LLMs are trained to simplify anything. They have billions of parameters after all.
- dTal 3y agoThey "simplify" the training data, which they are vastly smaller than. LLMs are like compression algorithms. You could imagine feeding the training data back in, letting it guess the next token, and entropy coding the residual - this would result in an excellent compression ratio. This compression performance is a direct consequence of abstract features of the dataset that it has managed to encode - knowing that the capital of France is Paris allows you to make predictions about many sentences, not just "The capital of France is...".
- mxkopy 3y agoTrue, but I still think there's some fallacy here. Are we sure that models of the world (i.e. understanding) are the only way to achieve compression?
- courseofaction 3y agoIntuitively, I think this also hints at why LLMs get more prone to confusion when trained to be "safe" - the underlying representations for applying human morality in context are much more complex to learn than simpler but potentially psychopathic logic.
- clarge1120 3y agoThis sounds correct. Humans are highly fickle and contradictory when it comes to morality. Even the Golden Rule is hotly contested. LLMs lose touch with reality as they try to navigate humanity’s moral landscape. Our current solution is to align an LLM to a worldview. The good news is that this will pit one LLM against others, and virtually eliminate any potential for a single powerful AI to emerge and do something harmful.
- freecodyx 3y agothe main thing about LLM's in my opinion is the tokenization part, words are already clustered and converted into numbers(vectors) it's already a big deal. we are using learned weights, the attention part feels like a brute force approach to learn how those vectors are likely used together (if you add positional encoding as an additional information). statistics on large amount of amount of data just seems to work after all.
- sanxiyn 3y agoThis is wrong, byte-level models work fine, even if not as well as word-level models. From comparison of byte-level models and word-level models, we know tokenization part is responsible for minuscule part of performance.
- kypro 3y agoI'm not sure I'm personally convinced LLMs are bad at arithmetic, I think they might just approach it differently to us. Something you'll find if you ever train a neural network to learn a mathematical function is that it will only ever approximate that function. It won't try to guess what the function is exactly like a human might do. For example consider, f(1) = 2, f(2) = 4, f(3) = 6, f(4) = 8, f(5) = 10. As a human you know how important precision is in maths and you know generally humans like round numbers so you naturally assume that, f(x) = x2 Neural networks don't have these biases by default. They'll look for a function that gets close enough maybe something like, f(x) = x1.993929910302942223 From a neural network's perspective the loss between this answer and the actual answer is almost so trivial that it's basically irrelevant. Then a human who likes round numbers comes along and asks the network, what's f(1,000)? To which the neural network replies, 19939.3 Then the human then goes away convinced the AI doesn't know maths, when in reality the AI basically does know maths, it just doesn't care as much about aromatic precession as the human does. Because again, to the AI 19939.3 is a perfectly acceptable answer. So now for fun let me ask ChatGPT some arithmetic questions... > ME > what's 2343423 + 9988733? > ChatGPT > The sum of 2343423 and 9988733 is 12392156. WRONG! It's actually 12332156. That's an entire digit out and almost 0.5% larger than the actual answer! > ME > what is 8379270 + 387299177? > ChatGPT > The sum of 8379270 and 387299177 is 395678447. Er, okay, that was right. Bad example, let me try again. > ME > what is 2233322223333 + 387299177? > ChatGPT > The sum of 2233322223333 and 387299177 is 2233322610510. WRONG! It's actually 2233709522510. That's 6 digits out and almost 0.02% smaller than the actual answer! If you take a more open minded view I think it's fair to say ChatGPT basically does know arithmetic, but its reward function probably didn't prioritise arithmetic precision in the same way a decade of schooling does for us humans. For ChatGPT having a few digits wrong in an arithmetic problem is probably less important that its reply containing that sum being slightly improperly worded. I guess what I'm saying is that I'm not sure I quite agree with the author that LLMs don't do arithmetic at all. It's not that they're trying to guess the next word without arithmetic, but more that they're not doing arithmetic the same as we humans do it. Which is may have been the point the author was making... I'm not really sure.
- SkyPuncher 3y agoLLMs are bad at math because they don't actually understand the rules of math. They can write code to do math, but without code they can only estimate how likely a series of numbers are to be seen together. They're very likely to get things like 2+2=4 correct because that's probably unique and common in their training data. They're unlikely to get two random numbers correct because it doesn't actually know what those numbers mean.
- golemotron 3y agoIn what world does "Quietly, quietly," have five syllables?
- pgspaintbrush 3y agoJapan =) Here's the original: https://basho-yamadera.com/en/yamadera/horohoro/ https://basho-yamadera.com/en/yamadera/horohoro/
- IIAOPSW 3y agoBecause the language we are teaching it is sophisticated enough to embed a Turing Machine?
- HALtheWise 3y agoI don't see enough discussion of the fact that LLMs are actually trained with two losses: text prediction and a regularization loss of some sort that effectively encourages the network to use "simple" internal structure. That means the training process isn't only trying to predict the next token, it's specifically trying to find the simplest explanation that predicts the next token. Given that the history of science is mostly driven by trying to find the simplest explanation for observed phenomenon, thinking about regularization makes it much less surprising that LLMs end up learning how the world "actually works".
- braindead_in 3y agoIt's so mind-boggling to think that our everyday reality can be encoded as weights and biases in a giant matrix. Maybe we are just weights and biases.
- sandsnuggler 3y agoWhy do people keep saying its good at math when we have no clue about the training data, and all they do is insert some examples in an unscientific way in a program we have no clue what's behind it or whether its one system or even multiple.