4 ms·
The one thing is that they seem to be using relatively small models. This may be a really damning result but I was under the impression that any generalization
by MeImCounting 3y ago
The one thing is that they seem to be using relatively small models. This may be a really damning result but I was under the impression that any generalization capabilities of LLMs appear in a non-linear fashion when you increase the parameter count to the tens of billions/trillions as in GPT4. It would be interesting if they could recreate the same experiment with a much larger model. Unfortunately I dont think thats likely to happen because of the resources required to train such models and the anti-open-source hysteria likely preventing larger models from being made publicly available much less the data they were trained on. Imagine that, stifling research and fearmongering reduces the usefulness of the science that does manage to get done.