3 ms·
That paper doesn't discuss GPT-4 at all. It does however contain this interesting excerpt (emphasis mine): > Although we may observe an emergent ability to occ
by feanaro 4y ago
That paper doesn't discuss GPT-4 at all. It does however contain this interesting excerpt (emphasis mine):
> Although we may observe an emergent ability to occur at a certain scale, it is possible that the ability could be later achieved at a smaller scale—in other words, model scale is not the singular factor for unlocking an emergent ability. As the science of training large language models progresses, certain abilities may be
unlocked for smaller models with new architectures, higher-quality data, or improved training procedures. For example, there are 14 BIG-Bench tasks5 for which LaMDA 137B and GPT-3 175B models perform at near-random, but PaLM 62B in fact achieves above-random performance, despite having fewer model parameters and training FLOPs.
So it's not obvious that it should be so straightforward.