3 ms·
I think this leaves we find that explore in language models regularized (or maybe augmented?) by a n-gram model: instead of predicing next token without any ext
by renonce 2y ago
I think this leaves we find that explore in language models regularized (or maybe augmented?) by a n-gram model: instead of predicing next token without any external knowledge, the n-gram predictions can be added to the softmax head as a “default” prediction. The language model’s job would then be to improve the accuracy on top of the n-gram prediction. This shifts some of the language model’s job to the n-gram predictor, which relies on traditional methods and not GPUs, saving a lot of computation.
EDIT: Oh so this thing takes 10TB disk space to keep an index while a LLM takes… 175GB (assuming GPT-3 in fp16). A huge resource requirement that cannot be ignored.