3 ms·
> Do you take issue with ... the idea that these were a gigantic leap? Yes, the gigantic leap was transformers. No one thought we peaked with 340M or 1.5B para
by haldujai 3y ago
> Do you take issue with ... the idea that these were a gigantic leap?
Yes, the gigantic leap was transformers. No one thought we peaked with 340M or 1.5B parameters, in fact expectations from early work was that massively scaling was going to achieve zero-shot capabilities rather the emergent capabilities you're alluding to which are essentially variations of in-context learning, a relative disappointment.
Subsequent improvements in GPT-4, which seem to be mostly just more RLHF and MoE, are similarly not surprising and are temporizing measures while hardware and datasets are limited. Jury is still out whether the billions spent are worth it in terms of actually getting us closer to AGI.
It seems to be worth it for OpenAI/MS who are trying to be first to market and establish vendor lock-in.