3 ms·
Yes, but k-token lookup was already a thing with markov chains. Transformers are indeed better, but just because they model language distributions better than
by wnoise 2y ago
Yes, but k-token lookup was already a thing with markov chains. Transformers are indeed better, but just because they model language distributions better than mostly-empty arrays of (token-count)^(context).