3 ms·
> Also I think it's important to mention that just a short while ago virtually no-one thought that shoving more layers into an llm would be enough to reach AGI.
by defgeneric 3y ago
> Also I think it's important to mention that just a short while ago virtually no-one thought that shoving more layers into an llm would be enough to reach AGI.
This was basically the strategy of the OpenAI team if I understand them correctly. Most researchers in the field looked down on LLMs and it was a big surprise when they turned out to perform so well. It also seems to be the reason the big players are playing catch up right now.
- whimsicalism 3y agoI think it was a surprise the behaviors that were unlocked at different perplexity levels, but I don't really agree that LLMs were "looked down on."
- TeMPOraL 3y agoMaybe not "looked down on", but more of "looked at as a promising avenue". I mean, 2-3 years ago, it felt LLMs are going to be nice storytellers at best. These days, we're wondering just how much of the overall process of "understanding" and "reasoning" can be reduced to adjacency search in sufficiently absurdly high-dimensional vector space.
- whimsicalism 3y agoPeople certainly knew that language modeling was a key unsupervised objective to unlock inference on language. I agree that I think they underestimated quite how useful a product could be built around just the language modeling objective, but it's still been critical for most NLP advances of the last ~6+ years.