2 ms·
I don't see how that makes a difference. Even if the models were 100x the size it's not like the same argument couldn't be carried through. There is no theoreti
by holonomically 5y ago
I don't see how that makes a difference. Even if the models were 100x the size it's not like the same argument couldn't be carried through. There is no theoretical reason to believe increasing the number of parameters does anything more than simply allow encoding more of the training set into the parameters. There seems to be a fundamental confusion about what these language models are actually doing, they're glorified compression algorithms. [1] There is no reason to expect any kind of generalization performance from them on common sense tasks.
1: https://bellard.org/nncp/ https://bellard.org/nncp/