4 ms·
hnfong was referring to emergent abilities, not to "the basics" – the definition of emergent abilities as they apply to language models is that they are not pre
by mpeg 2y ago
hnfong was referring to emergent abilities, not to "the basics" – the definition of emergent abilities as they apply to language models is that they are not present in smaller models, only in large ones.
That also implicitly defines that a model without "emergence" would not be large.
If you happen to know exactly why this happens, I'll definitely read your paper!
- _akhe 2y agoPrepare to be underwhelmed - no need for a paper for this one: The only emergent ability of an LLM is that the model is more robust when there are more samples. When the number of samples is in the trillions instead of thousands, there are a lot more complete concepts available for it to match to your query. The so-called "emergent ability" of a chat-focused LLM is its accuracy, which is only possible with enough sample data, and also only possible if it's good at matching a query to the right samples, and wording a response in a way that's pleasing to the end user - something even the mainstream LLMs struggle to do well. IME, pretty much only GPT-4 and Mistral are even that good. Most of them aren't that good at anything yet, and half the time don't follow what I ask at all. Being a very large model is only part of what makes it good. The sad state of Hacker News is not how many people conflate "LLM" with "chatbot" but that an academic paper holds more weight than a library you can npm install right now and try. Seems we lost our way if abstract appeals to authority hold more weight than evidence before your eyes. Thanks for your comment! I do appreciate the discussion.