2 ms·
Please note that these results were obtained using a small amount of compute (compared to say a large language model training run) on a limited training set. No
by iandanforth 3y ago
Please note that these results were obtained using a small amount of compute (compared to say a large language model training run) on a limited training set. Nothing in the paper makes me think that this won't scale. I wouldn't be surprised to see a AAA quality version of this within a few months.