4 ms·
Let's break this down. Large Language Models are almost all transformer models. Transformer models are sequence prediction models. There are other kinds of sequ
by iandanforth 3y ago
Let's break this down. Large Language Models are almost all transformer models. Transformer models are sequence prediction models. There are other kinds of sequence prediction models like RNNs. Transformers are not tied to tokens, neither are RNNs. You can use them to predict sequences of values directly, you don't need to have high dimensional representations as input.
Now, why do LLMs work? They are exploiting structure. Language has structure, and language represents other structures. At a large enough scale the language models can learn and use that structure. There's quite a bit of consistency in language. You can train a model to be fluent (grammatically correct) without massive scale, but output often lacks meaning or coherency. It takes a lot of training to learn that "reading" a book and "reading" from a spinner hard drive platter are conceptually similar and extremely different in their details.
So, can you use a transformer with raw numerical data? Yes. Can you train a very large model on a very large amount of numerical data. Yes. Would you expect that training across a mixed corpus of raw numerical data would lead to a similar kind of 'understanding' that we've seen in LLMs? No. but it is not impossible.
Recall that language is just a layer on top of sensory processing. It helps a great deal but there is plenty of "intelligence" in the animal kingdom without it.
My personal opinion is that helping models build their own internal representations via curriculum learning and making sure it is being sent data which is generally related to, or correlated with, other data is very useful. If you don't have the aid of a universal sematic representation scheme like language, why not try your best to make it easy to make one?