14 ms·
Are emergent abilities of large language models a mirage?
- raydiatian 3y ago“Emergent anything” is probably the most obnoxious buzzword in all of machine learning.
- joaquincabezas 3y agoall you need is emergent anything
- raydiatian 3y agoEmergency emergent emergence
- thomastjeffery 3y agoIt gives finality to the idea that we don't understand the thing or where it came from. Why? Did we just lose all interest in understanding things? Wasn't that the whole point in the first place? Somehow, people are throwing up their hands and giving up at understanding the thing; yet at the same time they are acting like the thing will magically evolve into their wildest dreams! The most fundamental feature of LLMs is that they cannot be literal. They can only infer, never define. Why is it that the people studying LLMs think they have to emulate that trait? It's like they are only allowing themselves to look at it as a mysterious black box: to infer its behavior from its results. Did they forget that they are the ones who wrote the damn thing?
- amw-zero 3y agoThe whole idea of machine learning / AI is to build functionality indirectly though, i.e. to build a system which evolves into another system over time. They are inherently meta-systems, so it does make sense to think of them differently.
- Mike_12345 3y ago> Somehow, people are throwing up their hands and giving up at understanding the thing Go tell that to the researchers who are working hard on studying emergent properties.
- thomastjeffery 3y agoYes: the effects of the thing, and not the thing itself.
- opportune 3y agoBefore blackbox AIs from deep learning there were basically a few different kinds of AI: one was basically “algorithm complicated enough that we thought it required intelligence”, another was “general problem solver” like you get by applying Constraint Satisfaction techniques and heuristics, and highly fine-tuned encodings of human knowledge and research (this is a decision tree clinicians use to perform a differential diagnosis of a fever, this is a function that finds edges in an image based on hand crafted CV algorithms). The first group is basically not AI, it was just assumed to be. The other two groups were fully explainable but required a ton of effort to get working outside of very tightly scoped situations. For a long time researchers thought that some combination of the two approaches would lead to more generalized models, but all attempts at morphing the two sucked ass because all knowledge had to be a hand crafted ontology of rules and atoms that could only explicitly encode relationships. Also, while computers can solve CSP/graph traversal algorithms impossibly fast compared to humans, those tasks are not good models for human cognition or tasks beyond stuff like crossword puzzles. You should consider that despite considerable effort, human brains are themselves black boxes. And you know less about your own knowledge than you think. I do not know where I learned that Timbuktu is both a placeholder name and real place, though I could go find evidence for both. I don’t have to expend any effort to distinguish the sounds of different words, I don’t know why two things can both taste “good”but in completely different ways. Nobody ever taught me that newly met acquaintances tended to not care to discuss current events in the business world, I just figured it out based on a collection of experiences whose individual instances I can’t even remember. Even the best neuroscientist could not tell you why neurons interacting in a certain way makes it so I can both drive a car and sing, or why one person’s brain seems to better at some or generalized tasks than another’s. And, well, deep learning overturned the paradigm of handcrafting AI systems by automating the process of “have the model produce this output from this input” without requiring humans to define the “how” beyond tuning the shape of the model, which was itself a hugely important innovation in reducing the human time required to build an AI system. But it’s not just faster to make these models, it’s so ridiculously better at making AI models for things like “is there a dog in this picture” that nobody would even consider doing those things without deep learning. You actually can fiddle with DNNs to get an idea of how they work similar to what we do with brains and CAT scans, you have it do some stuff with commonalities and you figure out which common parts get activated. This is easy to do with convolutional layers as they very commonly learn for themselves how to perform edge detection. Anyway, long story short, fully explainable AI utterly sucks ass at many tasks that are like a walk on the park for blackbox AI. And we cannot explain our own intelligence and knowledge except in terms of emergent phenomena, nor can we give the full provenance of some factoid or skill we have on demand (just like an LLM cannot tell you where it learned something) in many cases[0], so it seems reasonable that we’d be in the same situation with AI. [0] The main difference is that we have memory of the various discrete experiences of our lives (which we can associate with some knowledge or skill), and there is no binary separation between “learning mode” and “doing mode” or “active memory” and “long term memory” for us like with AI. We can definitely associate some knowledge with a particular event, but this seems like it could be a false ontological representation of our knowledge because if the knowledge and event were unimportant (like what you had for breakfast on a particular day) we’d forget both of them; it’s actually all the subsequent cases in which the knowledge and memory of the event came in handy that contribute to us being able to explain it.
- opportune 3y agoI’m sorry you find it obnoxious but emergent phenomena are everywhere in math and science and as annoying as it is to you, it also happens with AI. The quest for more generalized models boils down to studying emergent behavior because we could never prescriptively define all the parameters/behaviors/requirements necessary for such a complex outcome. We don’t even understand how the relatively easily observed interactions between neurons in our own brains result in emergent intelligence. What’s so impressive about LLMs is they understand the semantics of some concepts so well that they can consistently produce higher quality outputs for tasks like “explain this complicated concept with a nursery song from the perspective of a pirate” than humans could, with approximately no instances of that task in their training data. That is emergent behavior and it’s a pretty big deal.
- raydiatian 3y agoI agree that emergent behaviors are real, and important. I am skeptical, although not completely unconvinced, that LLMs like GPT are going to produce truly emergent phenomena, such as true first-principles logical reasoning. The limitations of the underlying transformer architecture itself are, in my opinion, the problem. The first problem is that the embedding space of the transformer needs to grow much, much larger, and it's already huge. This matters because you need to model the order of neurons in the brain. The second problem is that you're never going to train an LLM (as they're designed today) that is going to produce a truly good 'emergent-phenomena' answer without multiple network traversals. This is because the human mind constantly and autonomously refines its thoughts. Perhaps a good counter-argument is that emergent phenomena are fundamentally a space-time domain concept. I am aware that things like Conway's Game of Life are a fantastic counterargument to my "the transformer architecture doesn't support it" argument. But I agree that the definition of "emergent behavior" when it comes to machine learning is too easily corrupted to be novel rather than rigorous.
- modeless 3y agoThe title of this paper is misleading. They are not arguing that the abilities are a mirage. They are arguing that the sudden ("emergent") appearance of unexpected abilities is not actually sudden, but gradual and predictable with model scale, if measured in an improved way.
- TheDudeMan 3y ago"Emergent" doesn't mean sudden. (That's not on you but on them.)
- Ygg2 3y agoIt does mean that certain properties not found in constituents is in the greater system (pilots in game of life can't do addition but can be used to make an adder that counts values).
- ianbutler 3y agoSo if I'm reading this halfway correctly, quality isn't suddenly emergent, it's continuous and gradual based on size of the model. It only appears emergent when researchers pick bad metrics. I and I assumed a lot of people, already thought performance was a function on model size (# of parameters). Is this not what the prevailing thought is for DNN performance? Agreed with the other posters that this title is misleading.
- freehorse 3y ago> I and I assumed a lot of people, already thought performance was a function on model size (# of parameters). I guess the disagreement has been in whether this function is "continuous" or not. I do not think the title is misleading, considering the article answers to quite specific claims in other articles. I agree it sounds misleading if you do not put it in that context.
- sgt101 3y agoI think most people expect(ed) that performance vs. size would (is) be an s-curve. The surprise for most is that we have climbed up the slopes so far and so fast. What the shape is is not clear to me.
- thomastjeffery 3y agoThe mistake is arriving these abilities to the model itself, and not the content being modeled. Text contains more data than language. Large Language Models work implicitly: they are not limited to finding language-specific patterns in the text that they model. Humans look at LLMs through a lens of expectation. Any time we find a feature we did not expect, we categorize it after-the-fact. That's our biggest mistake: LLMs are not made of categories!
- ChatGTP 3y agoThis is a very interesting way of looking at it...
- 6gvONxR4sf7o 3y agoThis is interesting. There's another implication here. That reliability/usefulness is an "emergent" phenomenon as underlying abilities become more accurate. It's the difference between siri not understanding you 1 word out of 10 (very accurate!), and it basically just understanding you. It's a continuous accuracy function and a discontinuous usefulness function.
- p-e-w 3y agoI'm quite skeptical of analyses like this one, because I doubt the metrics themselves. Emergence is something that is intuitively noticed by human observers. The desire to quantify everything then leads to the creation of (imperfect) metrics designed to capture what the observers already know. Those same metrics are then taken as the definition of the properties said to be emergent, and articles like this one are among the consequences of that choice. The paper's claim is essentially "these metrics which appear to demonstrate emergence can be replaced by other metrics that also represent model behavior, but that do not have scale discontinuities, so emergence isn't a real phenomenon". But an equally valid interpretation would be "none of these metrics actually capture the properties we are truly interested in". Which, given the complexity of what we are dealing with here, seems entirely reasonable. It's not like we suddenly learned how to accurately quantify performance at language tasks. The whole reason LLMs are so great in the first place is because traditional 'mechanical' language models suck so bad.
- amw-zero 3y agoYou just described science, and why you don't believe in science.
- p-e-w 3y agoIf by "don't believe in science" you mean "don't believe that every metric claimed to be representing a phenomenon actually does represent that phenomenon", you are correct.
- Centigonal 3y agoI think the claim that "these metrics which appear to demonstrate emergence can be replaced by other metrics that also represent model behavior, but that do not have scale discontinuities (emergence)" is really powerful. That might mean that behaviors we consider emergent are the consequence of a process that scales continuously with model size. i.e.: there may exist a bijection between, say, a step-function `can_do_arithmetic(size)` and a smooth, continuous function `arithmetic_skill_metric(size)` If we can use continuous metrics to back out the step-function equivalents, that'll help us predict when and how to get particular abilities to "emerge." For example: If a change results in a steeper slope on the continuous metric, we can predict it would cause the associated capability to emerge at relatively smaller model sizes.
- _8j50 3y agoI had someone much knowledgable on this topic than myself claim ChatGPT and the like "understand" stuff. My standing argument is that they wouldn't hallucinate incorrect responses if they understood anything, the hallucination is when their approximation of what a real response would be falls short. These emergent abilities are not actually that, but a result of humans' poor understanding of cognition and communication. What concerns me very much is how the harms that can be caused by LLMs has been so greatly under reported. I imagine being someone with the right power and access telling ChatGPT "find all people that would vote against this candidate in real time and devise ad content and social media messaging and bit interaction to change their minds or discourage them from voting" heck, any intel org of a major country is probably already working on this. No more whistleblowing or posting anonymously on social media, companies would even share models based on private email and conversations you had so other companies could use LLMs to identify everything you posted elsewhere and to have LLMs designate a score for hoe hireable you are. Police can crack down on crime better but also crack down on dissent or any police reforms. And we aren't even talking about war time use of LLMs or what happens when you marry something like ChatGPT with Dall-E and make it all real-time. I am warning anyone who will listen. Smartphones are the most dangerous things out there. Any service or interaction that depends on them is deteimental to peace and liberty of the masses long term. People have not learned a thing from Snowden or 2016 elections. And why are all the smart journalists asleep on the job on this topic. Where are the unreasonable scaremongerers when you need them!
- Mike_12345 3y ago> I had someone much knowledgable on this topic than myself claim ChatGPT and the like "understand" stuff. My standing argument is that they wouldn't hallucinate incorrect responses if they understood anything, the hallucination is when their approximation of what a real response would be falls short. ChatGPT models semantic relationships in the data. That's what your smart buddy means by "understanding". That is a high dimensional model of the data set which infers abstract semantic relationships / concepts. But he would not claim that those semantic relationships are exactly the same as any human interpretation of the data (which you refer to as the "real" response). Language models also have limited reasoning abilities. They are capable of misunderstanding as much as understanding.
- colordrops 3y agoAny phenomenon that is not a fundamental property of reality is a mirage, or rather a fuzzy human construct on top of a conglomeration of phenomena without discrete boundaries. And even those "fundamental" properties are suspect.
- cjbprime 3y agoAs others have said, it's an awful title. Could instead be something like "Is the emergence aspect of Emergent Abilities in Large Language Models a Mirage?". Like, there's supposed to be nothing academic researchers like more than re-using the same word in a title, or making it into a clever pun or quip -- it's like the Dad Jokiest subfield -- but instead we just get a title that implies one common argument that people make, and actually delivers an unpredictable different argument that seems plausible but not necessarily interesting.
- jaidhyani 3y agoCould have gone with "More Comprehensive Metrics Are All You Need"
- mirekrusin 3y ago"Emergent Abilities in LLMs are not spontaneous"
- stilist 3y agoI have zero technical understanding of the math or statistics, but looking at the graphs it seems suspicious that supposed jumps happen across unrelated tasks and models at the same scales--for example, in figure 1, the discontinuities are consistently in the 10^22 to 10^24 range. Obviously I'm just going by what the authors have chosen to include, but I'd expect more variation. At best I'd assume it's something about LLMs in general.
- reubenmorais 3y agoThe number of data points is tiny. There's only a handful of LLMs trained from scratch in the world, and sizes of models released in a "generation" tend to be close to each other somewhat. The field is very open source so people all over are building on top of the same shared literature. Plus I'm sure there are leaks very often and companies then rush to train their own pet architecture to whatever parameter size the competition is about to release.
- kurthr 3y agoI think that's just because there are only 2-3 points between 10^22 and 10^24, which is more about the data available (and that they have just seen dramatic improvements) than the measures or models themselves.
- mercer 3y agoCould that be something to do with the things I keep reading about how somehow knowledge from, say, an LLM for generative text somehow carries over (in some way) to an LLM for image generation? I'm obviously not very knowledgeable in this area :).
- derrickrburns 3y agoHere is an analogy. Two softball players. One can hit the ball an average of 230 feet, 40% of the at bats. The other can hit the ball an average of 210 feet, 40% of the at bats. The homerun wall is 220 feet. One is a GREAT homerun hitter. The other has a poor batting average. The issue is that the success measure is non-linear.
- derrickrburns 3y agoHere is an analogy. Two softball players. One hits the ball 230 feet on average. The other hits the ball 210 feet on average. The homerun fence is at 220 feet. One is considered a GREAT homerun hitter. The other is considered a poor one. The measure is non-linear. That takes nothing away from the GREAT homerun hitter.
- deleted 3y ago[deleted]
- m3kw9 3y agoIs it that they don’t understand how the models derive outputs like “step by step reasoning”, and then say this is an emergent behaviour?
- pcrh 3y agoI'm not a mathematician, but it appears to me that "emergent" properties are being defined as those which do not appear in a minor form below a threshold. However, many natural phenomena that are fully explainable from first principles show this property, giving rise to sigmoidal "S-curves", as shown in Figure 1.
- usgroup 3y agoI think that part of the reasons conclusions about emergence are tenable is due to the opaque nature of transformer architectures. For example if it was possible to train a Hidden Markov Model with billions of hidden states on a trillion tokens, you could more literally look and see what was going on. Other than not being able to scale HMMs to this kind of scale, is there any good reason to believe they would not perform equally well but without the magic?
- kolinko 3y agoYou're using billions and trillions loosely here. The hidden state in HMM would be (num_tokens ^ context size), so something like 60000^2000.
- usgroup 3y agoI'm not sure that calculation is correct, but say it is, perhaps a Variable Length Markov Chain then.
- kolinko 3y agoVariable length markov chains would merge some states, sure, but it will still be a similar order of magnitudes. Anything longer than 4 tokens/words of context and you bump into 30k ^ 4-10 -> cross a billion/trillion state boundary and you lose any chance of using markov chains. Also - but here I may be wrong - there is no way to "train" markov chains to do generalisations - that is, if a given sentence didn't appear on the internet, it won't be available as a state for the chain. In this aspect they are more similar to a database than anything else.
- seydor 3y agoPeople are right to doubt the claims of notOpenAI and others about the capabilities of their models. The nonlinear output gains do not mean that the quest for intelligence is over. It's already hard to steer them with RL to make proper math. It's more likely that the transformer will only be a part of the larger architecture.
- Animats 3y agoOK, the system improves with scale. For some metrics which have thresholds of success, that looks like a discontinuity. But the discontinuity comes from the metric, not the improvement. Anything measured by "winning" has this property. Small changes near the "winning" threshold result in large changes in wins. This is well known in sports. Is there more to this issue than this amplification effect?
- nintendo1889 3y agoreplying to your old message [1]. The openqnx monartis source code is on github: https://github.com/vocho/openqnx https://github.com/vocho/openqnx [1] https://news.ycombinator.com/item?id=26255095 https://news.ycombinator.com/item?id=26255095
- tunesmith 3y agoI've always taken emergence as just a word from the perspective of the beholder. It isn't anything essential to the thing itself. If you understand a complex system enough, emergence goes away and it's reductive again. But that's not to say that emergence as a concept isn't useful. It's very much about our relationship to our discoveries and how much we understand them.
- deleted 3y ago[deleted]
- Gordonjcp 3y agoAren't LLMs basically just Eliza with a huge a priori dataset?
- theonlybutlet 3y agoIs human consciousness a mirage? (I'd say yes, a complex arrangement much more simple things).
- 6451937099 3y ago[dead]
- 6451937099 3y ago[dead]