3 ms·
Our results demonstrate that the power of language models can be attributed, to a great extent, to the auto-regressive next-token training scheme, and not neces
by version_five 3y ago
Our results demonstrate that the power of language models can be attributed, to a great extent, to the auto-regressive next-token training scheme, and not necessarily to a particular choice of architecture.
I think that's obvious isn't it? Neural networks are universal function approximates, the question is how to make the efficient, either in parameters or computation or whatever, as well as all the usual stuff like encouraging convergence, avoiding big gradients, etc. That's why transformers are popular, nobody thinks they especially can compute a function that other models can't.
- nonrandomstring 3y agoYeah I immediately thought, isn't this congruent to a statement about Turing machines. Sure there are classes of many things that are computationally equivalent, including computers made of paper-tape, tin cans and string. Just most of them are horrendously inefficient and useful only as thought experiments. I saw this again in audio synthesis but with more nuance. Most methods are equivalent in some crazy limit, but all have a "special" area of most useful effectiveness. For example in theory you can predict a signal of many minutes or hours just using linear prediction (LPC), but only at the cost of a gargantuan parameter space that's less efficient than just sampling the signal. Nonetheless it is nice to see that researchers are connecting up these dots, even if the pure maths behind it isn't saying anything obviously useful right away. Who knows what insights this might lead to for discovering other new methods of computation.