3 ms·
Yet there are emergent behaviours from these LLMs that are both surprising and not immediately understood. [1][2][3] Everyone has theories, of course, but stil
by celestialcheese 3y ago
Yet there are emergent behaviours from these LLMs that are both surprising and not immediately understood. [1][2][3] Everyone has theories, of course, but still pretty "magic" considering these behaviours weren't theorised in papers prior to observation.
1 - https://www.jasonwei.net/blog/emergence https://www.jasonwei.net/blog/emergence
2 - https://arxiv.org/pdf/2206.07682.pdf https://arxiv.org/pdf/2206.07682.pdf
3 - https://www.quantamagazine.org/the-unpredictable-abilities-emerging-from-large-ai-models-20230316/ https://www.quantamagazine.org/the-unpredictable-abilities-e...
- kerkeslager 3y agoDon't cite stuff you didn't read or understand. [1] Is a summary of [2], by one of its authors, not a separate source. [2] Defines "emergent behaviors" in a way that you're clearly misunderstanding (because "emergent behaviors" is an extraordinarily poor way of communicating this--it's partly the fault of the researchers who chose this ambiguous language). All it's saying is that bigger models can do things that smaller models can't, which should be surprising to no one. It's NOT saying that the capabilities are anything more than the sum of the input data. [3] Is written by a journalist, not an AI researcher, and so it's limited by the things the journalist is excited about. The journalist, for example, downplays sections like, "The other, less sensational possibility, she said, is that what appears to be emergent may instead be the culmination of an internal, statistics-driven process that works through chain-of-thought-type reasoning. Large LLMs may simply be learning heuristics that are out of reach for those with fewer parameters or lower-quality data." If you're going to try to gather things from journalists rather than subject matter experts, you need to understand how journalists work, and how subject matter experts work, and look for paragraphs like that to understand what's actually happening.
- celestialcheese 3y ago> [1] Is a summary of [2], by one of its authors, not a separate source. Yes. Your point? I included both because I found them both interesting. The paper is the source, the 137 emergent behaviours page is one of the authors continuing the work, and [3] is a journalist talking about this, so I included it as it's a unique perspective. I used the word "emergent" because that's what the SME used when describing this. From 5.1 in the paper linked: > Although there are dozens of examples of emergent abilities, there are currently few compelling explanations for why such abilities emerge in the way they do. You say this "should be surprising to no one", yet the authors disagree. Additionally, in the GPT-4 system card - "Emergent" appears 15 times, specificly section 2.9 is interesting https://cdn.openai.com/papers/gpt-4-system-card.pdf https://cdn.openai.com/papers/gpt-4-system-card.pdf So it's not just a word used callously by one group of researchers at Google.
- deleted 3y ago[deleted]
- YeGoblynQueenne 3y agoNot the OP but the point they're trying to make (not very politely) is that you're largely uncritically repeating claims made by OpenAI and Google about their commercial products. If you do that, you're not helping anyone, including yourself, understand what's really going on. Where I went to school, the teachers kept repeating that they wanted us to think critically about stuff when we were writing essays for homework. They didn't really know how to teach that, so most kids didn't learn it, but it's a valuable life skill: learn to think about what people say, and why, not just take everything everyone says at face value, especially when the person is in a position of some sort of authority (like being the people who released a system, or who wrote a paper, or, hashem yerachem, The Godfathers of AI). Authority is the mind killer.