3 ms·
Right, for a human, presumably it would be easier to go ahead and learn the language. But that begs the question. We already know LLMs tokenize words, so they
by pwinnski 4y ago
Right, for a human, presumably it would be easier to go ahead and learn the language. But that begs the question.
We already know LLMs tokenize words, so they have definitely generated and stored such a table. The marvelous step forward from previous word-based solutions is that they store not just single words, not just pairs of words, but larger sequences of words, even up to short sentence length. It makes the table pretty large, as you suggest, but that's exactly we've been seeing, and each generation of GPT is that much larger than the one before.